A primer design method for multiplex PCR targeted sequencing technology
By building a high-quality genome library and using phylogenetic tree clustering to optimize primer design, the genome library quality screening, primer compatibility and dimer problems in multiple PCR technology are solved, and the detection efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202411805012.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-10
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2044-12-10
AI Technical Summary
Existing multiplex PCR technology has challenges in genome library quality screening, primer compatibility and primer dimer issues, resulting in inefficient detection and insufficient accuracy.
By building a high-quality genome library, clustering genomes using phylogenetic tree and designing primers, combining primer coverage, specificity, and dimer binding scores, primer design is optimized for compatibility and efficiency.
Improve the efficiency and accuracy of multiple PCR detection, reduce repetitive work, ensure high compatibility and coverage of primers in diverse genomes, and reduce the impact of dimer binding.
Smart Images

Figure CN119296644B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bioinformatics, molecular biology, microbiology and sequencing technology, and in particular to a primer design method for multiplex PCR targeted sequencing technology. Background Art
[0002] With the development of science and technology, significant progress has been made in the fields of molecular biology, genetics, bioinformatics and sequencing technology. Among them, polymerase chain reaction (PCR), as one of the cornerstone technologies of molecular biology, amplifies specific DNA fragments in vitro, so that trace amounts of DNA samples can be detected; while second- and third-generation sequencing technologies can sequence a large number of DNA fragments in parallel, thereby realizing rapid sequencing of the entire genome or transcriptome. Multiplex polymerase chain reaction (mPCR) technology is an extension of PCR technology, which allows multiple pairs of primers to be used simultaneously in a PCR system for simultaneous amplification of multiple target sequences. Compared with traditional PCR, this method has the advantages of high efficiency, high throughput and low cost, and can meet the clinical needs for simultaneous detection of multiple pathogens and microorganisms, thereby greatly improving diagnostic efficiency and accuracy. However, in practical applications, these technologies still have some challenges and limitations, which are mainly manifested in the following points:
[0003] 1. Genomic library quality screening: The construction of a high-quality species genome library is the foundation for the subsequent selection of species-specific segments and the design of ideal primers. Existing genome screening methods often rely on genome assembly-level tags (such as representative genome, reference genome, complete genome, and chromosome), lacking more detailed screening criteria (such as genome sequencing time, host, origin, and completeness), resulting in uneven genome quality in genomic libraries.
[0004] 2. Primer compatibility: When designing multiplex PCR primers to detect target pathogens, the primers themselves need to have high compatibility, as the genomes of target pathogens in clinical samples are diverse. That is, the primers for each target pathogen must be able to copy the diverse genomes with slightly different sequences within that pathogen. Existing methods typically first search for highly conserved species-specific fragments within the species genome, and then use these specific fragments as templates to design primers. However, for species with more diverse genomes, the process of finding specific fragments is often not smooth, requiring continuous adjustment of the genome library for multiple rounds of specific sequence searches. This workload is not only large, but also difficult to obtain specific fragments with a length that meets the primer design requirements and high coverage.
[0005] 3. Primer-dimer issues: Due to the large number of primers in multiplex PCR systems, primer-dimer binding can reduce system efficiency. Common methods screen and optimize primer-dimers based solely on binding free energy values, without fully considering the impact of factors such as the base length of the primer-dimer binding region and the distance between the start and end points of the dimer binding region and the 3' ends of the two primers on dimer binding. Summary of the Invention
[0006] Purpose of the invention: The purpose of the present invention is to provide a primer design method for multiplex PCR targeted sequencing of pathogens, so as to avoid a lot of duplication of work in obtaining species-specific sequences with diverse intraspecies genomes and provide ideal compatibility.
[0007] Technical solution: The present invention provides a primer design method for multiplex PCR targeted sequencing technology, which comprises the following steps in sequence:
[0008] S100: Determine the optimal reference genome and construct a high-quality genome library of the target pathogen;
[0009] S200: constructing a phylogenetic tree for the high-quality genomic library;
[0010] S300: traversing the optimal reference genome with a fixed step size to obtain candidate primer pairs;
[0011] S400: Evaluate the candidate primer pairs, including scoring of coverage, specificity, and dimer binding.
[0012] The target pathogens described in the present invention refer to a variety of organisms targeted by pre-designed primers in a targeted multiplex PCR primer system, including but not limited to viruses, bacteria and fungi.
[0013] The primer design and optimization method provided by the present invention is applicable to the field of high-throughput molecular diagnosis, and is particularly suitable for multiplex PCR to detect multiple target pathogen targets through a single PCR reaction. Due to the diversity of target pathogen genomes in clinical samples, it is particularly important to select primers with high compatibility to ensure that diverse genomes with slight sequence differences can be effectively copied. Building a high-quality and diverse species genome library is the basis for ensuring high compatibility of primers. In this way, while ensuring that the primers highly cover the genome library, it can be inferred that the primers also have high compatibility with various genomes in the real world. At the same time, when designing primers, it is also necessary to improve the compatibility of primers to diverse genomes and adopt an efficient design strategy to reduce tedious workload. The present invention can solve these two difficulties very well.
[0014] The present invention builds the high-quality genome library of target pathogen, is different from the prior art that screens according to the order of genome assembly level, takes into account the multiple measurement standards such as genome sequencing time, host, place of origin, integrity, so when screening, can filter out the genome of sequencing quality that is general because of sequencing year too early, sequencing technology is not enough, or the genome of host and place of origin and Chinese population too far from source, if these genomes are present in the genome library, can interfere with the assessment of primer compatibility.Meanwhile, in order to ensure that the genome library has diversity, the present invention also uses the method for building phylogenetic tree to carry out the clustering on kinship to the genome library, ensures that the genome in the library neither has most of the focus on a certain subspecies branch, poor diversity, nor the evolutionary branch of individual far source is omitted when designing primers.The construction of high-quality genome library is based on the best reference genome, obtains other same species genomes according to assembly level, genome sequencing time, host, place of origin, integrity etc., and requires the sequencing time of the same species genome to be close to the current year, and has a certain sample size.
[0015] In the primer design stage, in order to improve the compatibility of primers to the diverse genomes of the target species, a common method in the industry is to first obtain the conserved fragments shared by multiple genomes in the genomic library, and then design primers on these conserved fragments. In order to obtain these highly conserved fragments, it is necessary to perform multiple rounds of genome sequence alignment and clustering to find the shared fragments. This process is cumbersome and time-consuming, especially for species with high genomic diversity. After a lot of work, it is still not possible to smoothly find a sequence with both high conservatism and high species specificity. The present invention first determines an optimal reference genome in the species genome library, and then clusters the diverse species genome library with a phylogenetic tree. If there is a cluster from a distant source, it is then grouped separately, and an optimal reference genome is also confirmed in the grouping. Design candidate primers on the optimal reference genome, and then calculate the coverage of these candidate primers in the genomic library. For the genome grouping from a distant source, design candidate primers on the reference genome in the grouping, and then calculate the genome coverage of the candidate primers in the grouping. This method of first grouping, then designing primers, and finally calculating primer coverage reduces the large amount of long fragment alignment work when obtaining common fragments, is more concise, and also ensures high primer coverage.
[0016] Furthermore, the step S100 specifically includes the following steps:
[0017] S110: Obtain the best reference genome of the target pathogen;
[0018] S120: Acquire several genomes of other species that were sequenced most recently, wherein the genomes of other species are acquired in descending order of assembly level;
[0019] S130: Screening other genomes of the same species with high completeness and meeting the sample quantity. When the sample quantity is insufficient, the sequencing time of step S120 is preferentially extended.
[0020] Among them, the optimal reference genome is determined based on multiple dimensions such as genome source, genome assembly level, sequencing time (year) and CheckM value to ensure that the screened genome is of high quality and representative of the species.
[0021] Specifically, the method for determining the optimal reference genome is as follows: screening an optimal reference genome from the RefSeq database or the GenBank database in the order of priority of Representative Genome>Reference Genome>Complete Genome>Chromosome>Scaffold>Contig;
[0022] When there are multiple genomes or genome structures at the same level, the genome or genome structure with the most recent sequencing year and the highest CheckM value is selected;
[0023] Among them, the sequencing year has a higher priority than the CheckM value; and if the sequencing years and CheckM values of multiple genomes or genome structures of the same level are the same, one is selected arbitrarily.
[0024] Specifically, species in the RefSeq database typically contain at least one representative genome or reference genome. If a species has more than one representative genome or reference genome, the genome with the most recent sequencing year and the highest CheckM value is selected as the optimal reference genome, with sequencing year taking precedence over CheckM value. If both sequencing year and CheckM value are the same, a random reference genome is selected. If a species does not have a representative genome or reference genome in the Refseq database, the optimal reference genome is identified in the GenBank database using the same method as described above. If neither a representative genome nor a reference genome is available in either the Refseq or GenBank databases, genome structure selection is performed in the descending order of priority: Complete Genome > Chromosome > Scaffold > Contig. If the number of genomes / genome structures at the same level is greater than one, the genome with the most recent sequencing year and the highest CheckM value is selected as the optimal reference genome, with sequencing year taking precedence over CheckM value. If both sequencing year and CheckM value are the same, a random reference genome is selected.
[0025] The CheckM value used in the present invention to confirm the optimal reference genome is an indicator for measuring genome completeness.
[0026] Among them, for pathogens that are taxonomically classified as "species", the number of samples is not less than 10; for pathogens that are taxonomically classified as "species complex" or "genus", the number of samples is 50.
[0027] It should be noted that step S120 prioritizes genome samples with a complete assembly level. If the number of samples is insufficient, the sequencing time in step S120 is increased to obtain more samples. If this still does not result in a sufficient number of samples, chromosome, scaffold, and contig genes are downloaded according to assembly level priority until the required number of samples is reached.
[0028] As a further optimization of the present invention, each target pathogen in the high-quality genome library was tested for completeness using CheckM2 software, and genomes with completeness greater than or equal to 95% were screened.
[0029] In particular, the genomes of some species are more complex and CheckM2 cannot be scored. The CheckM items of NCBIRefSeq / GenBank are also 0. In this case, this indicator can be ignored.
[0030] The phylogenetic tree described in this application, also known as a molecular evolutionary tree, is a visual tree diagram constructed by comparing the numerical values of sequence differences among biological macromolecules (such as proteins and nucleic acids). It depicts the sequence-level relationships between different entities and is a bioinformatics analysis method. Prior art similar to the subject matter of this application rarely discloses the construction of phylogenetic trees to examine the evolutionary relationships of individual genomes within a species' high-quality genomes, and its application in primer design methods for multiplex PCR targeted sequencing technology. The present invention uses phylogenetic trees to cluster genomes based on genomic sequence differences, thereby reflecting the relationships between genomes. Specifically, the present invention uses Parsnp software to construct a phylogenetic tree. It first aligns the most similar sequence portions of each genome within a species to form a "core alignment region." Single nucleotide polymorphisms (SNPs) are then detected within this "core alignment region." These SNPs are generally considered the most reliable mutations for studying evolutionary relationships. A maximum likelihood algorithm is then used to construct an evolutionary tree based on the differences within the "core alignment region," reflecting the relationships between these genomes.
[0031] Furthermore, the genome evolutionary tree obtained in step S200 meets the following requirements: the branches of the genome evolutionary tree are uniform, there are no isolated branches with obviously distant evolutionary distances, and there is no situation where the vast majority of genomes are densely distributed on the same branch. In a phylogenetic tree, a leaf node represents the terminal node of the phylogenetic tree, and there are no further branches at this node. In this study, each leaf node represents a genome participating in the construction of the evolutionary tree. The initial evolutionary tree first analyzes the similarities and differences of multiple genomes with similar sequences to form an independent small branch structure, and multiple small branch structures then form branches extending toward the root node. The branch length represents the kinship between two nodes. The longer the branch, the more distant the kinship. The method for determining whether the evolutionary tree branches are uniform is: in the phylogenetic tree, in the small branch structure with the most leaf nodes, the number of leaf nodes is less than 30% to 50% of the total number of genomes. Furthermore, if the ratio of the length of certain evolutionary branches to the average length of all evolutionary branches is greater than 10-100 times, the genomes on these branches are considered distantly related. If the number of genomes on a branch is greater than 0.01-10% of the total number of genomes and the quality of these genomes is good, these genomes are grouped according to the evolutionary tree branches and primer design is subsequently performed within the group. If the number of genomes on a branch is greater than 0.01-10% of the total number of genomes, these genomes are directly deleted.
[0032] As a preferred embodiment of the present invention, whether genomes of obviously distant isolated branches are eliminated from the high-quality genome library, and whether or not genomes with an evolutionary distance that exceeds a certain proportion of the total number of genomes are grouped according to the evolutionary tree branches is evaluated based on 5% of the total number of genomes. It should be understood that in the design of multiplex PCR primers, in order to cope with the diversity of target pathogen genomes in clinical samples, it is necessary to ensure that the primers have high compatibility and can cover the diverse genome sequences of each target pathogen. Therefore, it is necessary to select genomes with high sequencing quality and universality to construct a high-quality genome library. This requires that the selected genomes should not be concentrated in a single evolutionary lineage (such as a certain subspecies) to avoid insufficient universality. Similarly, when a group of genomes is too distant in evolutionary relationship and the number exceeds 5% of the total, it should be divided into several groups according to the evolutionary relationship, and primers should be designed to cover them separately.
[0033] Regarding step S300, in step S300, shingled slicing of the whole genome is performed with a sliding window length of 300 bp and a sliding step length of 20 bp, and primers are designed on each fragment using primer3 software.
[0034] Preferably, for the second generation sequencing technology, the amplicon length is 180-300 bp, T m The value range is 55-65℃, the primer arm length is 18-25bp, and the primer arm GC base content is 40-60%.
[0035] For the third-generation sequencing technology, those skilled in the art may make adaptive changes to the design of the sliding window length and the sliding step size based on the above-mentioned second-generation sequencing technology without departing from the principles of the invention of this application, and all of these changes are within the scope of protection of this application.
[0036] For step S400, the coverage evaluation method is to match the candidate primer pairs to each other genome of the same species in the high-quality genome library and calculate the primer coverage:
[0037] Primer coverage = number of qualified matching genomes / total number of genomes × 100%.
[0038] Among them, a qualified match must at least meet the following conditions simultaneously: both the upstream primer and the downstream primer can achieve a match, the matching length of the primer amplicon is not less than 95% of the total amplicon length, and the primer amplicon length is 100-300bp.
[0039] The specific evaluation method comprises the following steps:
[0040] S421: aligning the candidate primer pairs in pairs in the NCBI nt library against all genomes in the library;
[0041] S422: After alignment, a primer insert is obtained. The primer insert refers to a nucleic acid sequence defined and amplified by a pair of complementary oligonucleotide primers during PCR. The starting end of the primer insert is defined by the 3' end of the forward primer, and the ending end is defined by the 5' end of the reverse primer.
[0042] S423: Inputting the primer inserts into Kraken 2 software for verification, thereby screening primer pairs with qualified specificity;
[0043] S424: When the species of the genome to which the primer insert fragment belongs is consistent with the primer target species, the species discrimination result of Kraken 2 must be that species; or when the species of the genome to which the primer insert fragment belongs is inconsistent with the primer target species, the species discrimination result of Kraken 2 must be consistent with the genome species.
[0044] This approach allows primers to bind to sequences from non-target species while ensuring that non-target species can be correctly identified based on primer inserts in Kraken 2. This relaxes the requirement for primer specificity and prevents false positives due to species misidentification, thus improving primer utilization.
[0045] In multiple-primer systems, one common factor that affects amplification efficiency is primer dimerization. This occurs when primers within a system bind to each other due to base complementarity, preventing them from binding to pathogen sequences and reducing amplification efficiency. A common approach in the industry is to first enumerate all possible dimerizations in the system and calculate their binding free energies. If the free energy falls outside a safe threshold (<-4 kcal / mol), the two primers are considered likely to form dimers in real-world experiments. Alternatively, a global optimization algorithm, such as a greedy algorithm, is used to combine primers, selecting new primers that minimize dimerization with existing primers. This cycle is repeated repeatedly until the desired primer count is reached. The calculation of dimer binding free energy involves base pairing, stacking interactions, ion concentration, and temperature. It simply measures the ease with which a segment of bases binds at a specific temperature and solution concentration, but overlooks the fact that dimer binding is also dependent on the distance between the dimer region and the 3' ends of the two primers. The closer the dimer binding region is to the 3' ends, the less favorable the binding effect. Global optimization algorithms often produce a fixed set of primer combinations, which can be challenging when replacing certain primers in the system, potentially requiring the algorithm to be rerun.
[0046] The present invention proposes a new primer-dimer evaluation method that not only takes into account the dimer binding free energy, which is of general concern to those skilled in the art, but also integrates factors such as the length of the primer-dimer binding region, the number of GC bases in the binding region, and the distance between the primer-dimer binding site and the 3' end of each primer. Traditional methods often rely on algorithms to evaluate primer pairs one by one in sequence, resulting in a lack of flexibility and controllability in the results. The evaluation method provided by the present invention is more intuitive and allows researchers to freely select primer pairs, rather than simply relying on fixed combinations automatically generated by an algorithm, thereby providing greater flexibility and freedom of choice.
[0047] Specifically, the dimer binding score is:
[0048] ;
[0049] When any one of the candidate primer pairs No less than 10 3 When , the primer pair corresponding to the primer is screened out;
[0050] The "dimer binding region" refers to the duplex region formed by partial or complete complementarity between two complementary oligonucleotide primers during the polymerase chain reaction (PCR) process. This region begins at the non-complementary region at the 3' end of one primer and may extend to the 3' end of the other primer or any other location, terminating at the end point of complementary pairing. Individual base mismatches are permitted within this region, but overall complementarity is sufficient to form a stable duplex structure. GC-count indicates the number of GC bases in the dimer binding region (mismatches are not counted). The first complementary base from the 5' end to the 3' end of one primer is counted as the start point, and the last complementary base is counted as the end point. A maximum of two base mismatches are tolerated in the start-end region. L represents the base length of the dimer binding region from start to end. d1 represents the distance from the start of the dimer binding region to the 3' end of the nearest primer, and d2 represents the distance from the end of the dimer binding region to the 3' end of the nearest primer.
[0051] The above indicators contribute to dimer evaluation, providing a more comprehensive assessment of the difficulty of dimer binding and the impact on the entire system. Ultimately, each candidate primer is assigned a score, representing the sum of all possible dimer scores generated by that primer and other primers in the system. Higher scores indicate more severe dimer binding. Researchers can directly select and combine primers based on these scores, providing greater flexibility than conventional evaluation algorithms.
[0052] Based on the dimer evaluation method provided by the present invention, when the total score of any primer in a pair of primers is greater than or equal to 10 3For example, the F-Primer and R-Primer shown below both have a dimer score of 4096 using the dimer score provided by the present invention, and should be evaluated as primer pairs that need to be eliminated.
[0053] Score: 6, Delta G = - 8.17 kcal / mol
[0054] F-Primer: CGTTCAGTACAATG CGGCCG
[0055] R-Primer: GCCGGC GTAACATGACTTGC
[0056] The underlined portion of the primer represents the dimer junction region, which has a binding length of six bases, all of which are GC bases. The dimer binding free energy is -8.17 kcal / mol, which is less than the threshold of -4 kcal / mol commonly used by those skilled in the art, indicating that dimers readily bind in the real world. This demonstrates that the dimer evaluation method provided by the present invention can objectively assess dimer binding and facilitate primer screening.
[0057] It should be noted that the technical solution claimed in this application uses the order of primer coverage, specificity, and dimer binding score as the preferred evaluation method, that is, primer pairs screened in the first evaluation are used in the subsequent evaluation. Nevertheless, those skilled in the art may freely change the evaluation order as necessary, and this is within the scope of protection of this application.
[0058] The primer design method provided by the present invention includes determining the optimal reference genome of the target pathogen, constructing a high-quality genome library of the pathogen, designing primers, and evaluating the coverage, specificity, and dimer binding of the primer pairs. The method first traverses the optimal reference genome of the target pathogen to design primer pairs. Then, the primers are aligned to the gene library to calculate coverage, thereby avoiding a lot of duplication of work when obtaining species-specific sequences with a relatively diverse genome within the species. The specificity of the candidate primer pairs is then tested to ensure that the fragments copied by the primer pairs in the real sample can correctly identify the species to which they belong. Finally, the candidate primer pairs are scored using a self-developed dimer scoring algorithm to reduce dimer binding in the primer system and improve the efficiency of the primer system. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 Flowchart for designing primer systems for multiplex PCR targeted pathogen detection;
[0060] Figure 2 This is the evolutionary tree of the high-quality genome of Staphylococcus worderi;
[0061] Figure 3 This is a phylogenetic tree of the high-quality genome of Candida tropicalis;
[0062] Figure 4 Phylogenetic tree of the high-quality measles virus genome. DETAILED DESCRIPTION
[0063] In order to make the technical solution of the present invention clearer, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0064] Example 1
[0065] This example uses Staphylococcus waltersii (bacteria), Candida tropicalis (fungus) and measles virus (virus) as target pathogens and provides a detailed description based on Example 1.
[0066] like Figure 1 As shown, the present invention provides a primer design method for multiplex PCR targeted sequencing technology, which comprises the following steps in sequence:
[0067] S100: Determine the optimal reference genome and construct a high-quality genome library of the target pathogen;
[0068] S200: constructing a phylogenetic tree for the high-quality genomic library;
[0069] S300: traversing the optimal reference genome with a fixed step size to obtain candidate primer pairs;
[0070] S400: Evaluate the candidate primer pairs, including scoring of coverage, specificity, and dimer binding.
[0071] S100. Construction of high-quality genomic libraries
[0072] S110. Obtaining the best reference genome for the target pathogen
[0073] Obtain the assembly_summary_refseq.txt file from the NCBI RefSeq database. Use the organism_name field and species_taxid field to identify the target species. Also, select the optimal reference genome based on whether the genome category field (refseq_category) is "reference" or "representative." Download the optimal reference genome for each target species.
[0074] The species name of Staphylococcus waltneri is Staphylococcus warneri , species ID: 1292; the species name of Candida tropicalis is Candida tropicalis , species ID: 5482; the species name of measles virus is Measles morbillivirus , species ID: 11234. The above three species can be directly confirmed from the RefSeq database and downloaded to the representative genome.
[0075] S120. Obtain the genomes of several other species with the most recent sequencing time
[0076] The target pathogen is determined based on the species name field (organism_name) obtained in step S110 and the species ID field (species_taxid); and other genomes of the same species whose sequencing information was published later than 2018 and whose host field (host) is "Homo sapiens" or "Human" are screened through the NCBI information page of the genome.
[0077] Among them, as a preferred embodiment, source fields (source) including "soil", "water" and other genomes that are not related to humans or mammals can be screened out. These genomic information is of little use in genomic research on human infectious pathogens.
[0078] Download other genomes of the same species in the order of assembly priority: complete genome > chromosome > scaffold > contig. When the target pathogen is at the "species" level, the total number of genomes must be no less than 10; when the target pathogen is at the "species complex" or "genus" level, the total number of genomes must be greater than or equal to 50.
[0079] If the number of genomes at the complete genome level for the current pathogen still does not meet the minimum number standard (10) after downloading all complete genome-level genomes, the sequencing time standard will be relaxed by 2 years (for example, the first relaxation of "2018 to present" to "2016 to present"). If the number of genomes still does not meet the standard, they will be downloaded according to the assembly level priority order until the number of genomes meets the standard.
[0080] S130. Screen for genomes of other species with high completeness and sufficient sample size
[0081] The best reference genomes of the three target pathogens were tested for genome integrity using CheckM2 software (CheckM2 v1.0.2). The test standard was that the score of the integrity index CheckM was not less than 95%.
[0082] If the number of unqualified genomes of the current pathogen is not less than 15%, then the candidate genomes are added and the completeness is evaluated according to S120. In this embodiment, because the species level of the three target pathogens is species, the total number of genomes of each species is required to be no less than 10. The genome of Staphylococcus worderi can be scored using CheckM2; the genomes of Candida tropicalis and measles virus cannot be scored using CheckM2, and the CheckM in the NCBI RefSeq and GenBank databases are both 0, so this indicator is not considered for these two pathogens. Table 1, Table 2, and Table 3 are the genomic information and CheckM scores of Staphylococcus worderi, Candida tropicalis, and measles virus, respectively.
[0083] Table 1 List of high-quality genome libraries of Staphylococcus waltneri
[0084]
[0085] Table 2 List of high-quality genome libraries of Candida tropicalis
[0086]
[0087] Table 3 List of high-quality measles virus genome libraries
[0088]
[0089] S200. Using Parsnp to construct a phylogenetic tree for high-quality genomic libraries
[0090] For each target pathogen screened genome, an evolutionary tree was constructed using Parsnp software (Parsnp v1.7.4), and then visualized using tools provided on the online website to test whether the genome was evenly distributed from the perspective of evolutionary relationships.
[0091] Using the best reference genome of each target pathogen as a reference, and all other genomes of the same species, a genome evolutionary tree is constructed. The branches of the evolutionary tree should be relatively uniform, with neither isolated branches with significantly long evolutionary distances nor all genomes densely concentrated on the same branch. Specific judgment method: In the small branch structure with the most leaf nodes, the number of leaf nodes should be less than 50% of the total number of genomes. For branches with a ratio of branch length to average branch length greater than 100, if the leaf nodes contained are greater than 10% of the total number of genomes and the quality of the genomes is good, these genomes are grouped according to the evolutionary tree branches, and primer design is subsequently performed within the group; if the number of leaf nodes contained is less than or equal to 10% of the total number of genomes, these genomes are directly deleted.
[0092] The phylogenetic trees of Staphylococcus waltneri, Candida tropicalis, and measles virus are shown in Figure 2 、 Figure 3 、 Figure 4 . Figure 2 In the figure, the dotted box circles the small branch structure with the most leaf nodes, with a total of 5 leaf nodes, which does not exceed 50% of the total genome. The average length of the evolutionary tree branches is 0.114, and the longest branch length is 0.389, which does not exceed 100 times the average length. The results show that the genome evolutionary tree is evenly distributed and there are no obvious isolated genomes, which is considered to be qualified for the genome library construction. Similarly, it can also be judged that Figure 3 、 Figure 4 The phylogenetic trees all met expectations. The high-quality genomic libraries constructed for the three species were qualified, and there was no need to delete genomes or group specific genomes for primer design. It should be noted that some branches in the phylogenetic tree were not included in the calculation of the longest branch length and average branch length because the software did not indicate branch length. These branches were not included in the calculation of the longest branch length and average branch length, nor were they deleted from the genotype library. However, they were included in the calculation of primer coverage of the genome in step S400, thus measuring the coverage of these genomes by the primers.
[0093] S300. Traverse the optimal reference genome with a fixed step size to obtain candidate primer pairs
[0094] Candidate primer pairs were designed iteratively on the best reference genome for each target pathogen.
[0095] Select the best reference genome in each target pathogen library. This example uses the second-generation sequencing application scenario as an example: with a sliding window length of 300 bp and a sliding step length of 20 bp, a shingled slice is performed on the whole genome, and primers are designed on each fragment using Primer3 software. If targeted amplification is performed on the second-generation sequencing platform, the primer design parameters are: the full length of the amplicon ranges from 180 to 300 bp; T m The value range is 55-65℃, where T m The values were predicted using the calculate_tm() function in the primer3 software. Primer arm lengths ranged from 18 to 25 bp. The GC content of the primer arms ranged from 40 to 60%, and the primer arms did not contain >= 5 consecutive identical bases. After design, identical primer pairs were removed.
[0096] When replacing with a third-generation sequencing platform, the present invention can match the requirements by changing the sliding window, sliding step and primer design parameters.
[0097] S400. Evaluate the candidate primer pairs, including coverage, specificity, and dimer binding scores
[0098] S410. Coverage evaluation
[0099] Calculate the coverage of each species' candidate primer pair in the high-quality genome library of that species. Match the candidate primer pairs obtained in step S300 to each other genome of the same species in the high-quality genome library, and calculate the primer coverage:
[0100] Primer coverage = number of qualified matching genomes / total number of genomes × 100%.
[0101] The criteria for a qualified match are: both the upstream and downstream primers can match the genome; the primer amplicon length is between 100-300bp; and the matching length of the amplicon is ≥ the total amplicon length × 95%; at least two consecutive bases from the 3' end of the primer arm completely match the genome.
[0102] After coverage evaluation, the primer information of the three target species that met various criteria is shown in Tables 3, 4, and 5. These candidate primers were used in the subsequent optimization steps.
[0103] Table 4 Primer information with qualified coverage for Staphylococcus worderi
[0104]
[0105] Table 5 Primer information with qualified coverage of Candida tropicalis
[0106]
[0107] Table 6 Primer information with qualified measles virus coverage
[0108]
[0109] S420. Specificity evaluation
[0110] Specificity evaluation includes the following steps:
[0111] S421: aligning the candidate primer pairs in pairs against all genomes in the NCBI nt library, i.e., primer-blasting;
[0112] S422: After alignment, a primer insert is obtained. The primer insert refers to a nucleic acid sequence defined and amplified by a pair of complementary oligonucleotide primers during PCR. The starting end of the primer insert is defined by the 3' end of the forward primer, and the ending end is defined by the 5' end of the reverse primer.
[0113] S423: Inputting the primer inserts into Kraken 2 software for verification, thereby screening primer pairs with qualified specificity;
[0114] S424: When the species of the genome to which the primer insert fragment belongs is consistent with the primer target species, the species discrimination result of Kraken 2 must be that species; or when the species of the genome to which the primer insert fragment belongs is inconsistent with the primer target species, the species discrimination result of Kraken 2 must be consistent with the genome species.
[0115] Species identification was performed using Kraken 2 (Kraken2 v2.1.3). Custom high-quality genomes of species must be added to the Kraken 2 library beforehand and mixed with the standard species library included with Kraken 2. The first 42bp of the candidate primer insert fragment was used for species identification in Kraken 2. If the species of the genome to which the insert fragment belongs is consistent with the primer target species or is a descendant of the target species, the species identification result in Kraken 2 must also be the target species or a descendant of the target species. If the species of the genome to which the insert fragment belongs is inconsistent with the primer target species and is not a descendant of the target species, the species identification result in Kraken 2 must be consistent with the genome species. This allows for the presence of copies of non-target species in the primer pair, but the species must be correctly identified in Kraken 2 to avoid false positives.
[0116] Tables 7-9 show the percentages of the four situations of “copying the target species and reporting correctly”, “copying the non-target species and reporting correctly”, “copying the target species and reporting incorrectly”, and “copying the non-target species and reporting incorrectly” after simulated amplification of the NCBI nt library and species identification by the candidate primers. Primers were considered qualified when the percentage of “copying the target species and reporting correctly” was ≥ 90%, and the percentages of “copying the target species and reporting incorrectly” and “copying the non-target species and reporting incorrectly” were both 0, or when there were special circumstances that could explain it.
[0117] Among the candidate primers for Staphylococcus vorticella in Table 7, SW_1_F / SW_1_R and SW_3_F / SW_3_R were considered to be unqualified primers with low efficiency because they "copied non-target species and reported incorrectly" and the ratio of "copied target species and reported correctly" was less than 90%. The other primers were qualified. All candidate primers for Candida tropicalis in Table 8 were qualified in specificity. It should be noted that among the candidate primers for measles virus in Table 9, although each pair of primers had a small proportion of "copied non-target species and reported incorrectly", the reason was that the artificial synthetic plasmid with taxid 32630 was amplified and reported as the target species, or the artificial synthetic vector with taxid 2877938 was amplified and reported as the target species. Such artificial synthetic species do not exist in real-world human samples, and therefore do not affect the amplification and species reporting of the primers. Therefore, all candidate primers are considered qualified.
[0118] Table 7 Specificity verification information of candidate primers for Staphylococcus worteri
[0119]
[0120] Table 8 Specificity verification information of candidate primers for Candida tropicalis
[0121]
[0122] Table 9 Specificity verification information of measles virus candidate primers
[0123]
[0124] S430. Dimer binding score
[0125] All qualified candidate primers in step S420 were enumerated using the software MFEprimer (MFEprimer v3.2.7) to enumerate all possible dimer binding situations. The dimer binding length must be greater than or equal to 3 bases and the binding region is allowed to have a mismatch of less than or equal to 2 bases.
[0126] On this basis, each dimer binding is scored, and the scoring rules are:
[0127] ;
[0128] Where GC-count represents the number of GC bases in the dimer binding region, L represents the base length of the dimer binding region, d1 represents the distance from the start point of the dimer binding region to the 3' end of the nearest primer, and d2 represents the distance from the end point of the dimer binding region to the 3' end of the nearest primer.
[0129] Scoring of a primer pair Equal to two primers All primer pairs are based on Values are sorted in ascending order, The smaller the value, the less dimer binding of the primer pair. In the multiplex PCR system, the present invention considers that the ≥10 3 is unqualified.
[0130] Tables 10, 11, and 12 show the dimer scores of all qualified candidate primers in step 7, among which CT_1_R and MM_2_F have dimer binding and the score is 8192 greater than 10. 3 , which is considered to be a relatively easy dimer to form. Therefore, the two primer pairs CT_1 and MM_2 were eliminated, and the remaining candidate primers were qualified. The dimer formed by CT_1 and MM_2 is shown below, where the two bases in brackets represent mismatches in the dimer binding region.
[0131] CT_1_R:GTAGAATAGTTAACT CGG(GA)TCTGT
[0132] MM_2_F: GCC(AG)AGACA GTAATTGGAGGGG
[0133] Table 10 Dimer verification information of candidate primers for Staphylococcus worteri
[0134]
[0135] Table 11 Dimer validation information of candidate primers for Candida tropicalis
[0136]
[0137] Table 12 Dimer verification information of candidate measles virus primers
[0138]
[0139] Example 2
[0140] In this example, the primers obtained by screening in Example 1 were subjected to a single-pair PCR wet test to verify the feasibility of the method described in Example 1.
[0141] (1) Test reagents and materials
[0142] Human cerebrospinal fluid sample; sample pathogens; self-developed nucleic acid extraction kit; reverse transcription kit (Abclonal, cat. no. RM20428); SYBR Green qPCR kit (Acori Biotech, cat. no. AG11701).
[0143] (2) Test instruments
[0144] TL2010S medium-throughput tissue grinder (Dinghaoyuan Technology, catalog number 0401265); automatic nucleic acid extractor (hema, catalog number E96-II); LightCycler 480Ⅱ real-time fluorescence quantitative PCR instrument (Roche, catalog number LC480); PCR instrument (Longji, catalog number A300); vortex mixer Vortex 5 (Qilin Bell, catalog number HD201805135); MiniL-12G large-capacity mini centrifuge (Maigao, catalog number MinilL-12G), etc.
[0145] (3) Primer synthesis and preparation
[0146] In single-plex PCR experiments, a universal sequence must be added to the specific primer sequence during primer synthesis. The upstream primer is composed of the universal sequence, UMI, and a specific upstream primer, and the downstream primer is composed of the universal sequence and a specific downstream primer. The primers are synthesized by Sangon Biotech Co., Ltd. and Aikerui Biotech Co., Ltd.
[0147] (4) PCR reaction system and reaction procedure
[0148] The amplification efficiency of the primers was evaluated using a QPCR system. Using pure bacteria or plasmids as templates, two groups, a positive template group and a blank control group, were set up. The QPCR reaction was performed according to the QPCR reaction system in Table 13 and the QPCR reaction procedure in Table 14. The amplification efficiency of a single pair of primers was evaluated based on the cycle threshold (Ct) and melting temperature (Tm value). In particular, because measles virus is an RNA virus, the sample used is RNA, which needs to be reverse transcribed before the QPCR reaction. The X in the table indicates that the specific input amount is determined according to the actual situation.
[0149] Table 13 PCR reaction system
[0150]
[0151] Table 14 PCR reaction program
[0152]
[0153] (5) Qualification criteria for single-plex PCR primers
[0154] The amplification efficiency of a single primer pair was evaluated using SYBR Green dye fluorescence quantitative PCR and classified as excellent, good, or poor. The evaluation factors were:
[0155] ① Blank template Ct-positive template Ct ≥ 5;
[0156] ② No amplification of blank template or blank template Ct ≥ 25;
[0157] ③ | Positive template T m Value - Blank Template T m Value|≥2.
[0158] Meeting all three of the above conditions is considered excellent, meeting at least one condition is considered good (condition ① must be met first), and meeting none of the three conditions is considered poor. Table 15 shows the single-plex PCR results for all primers.
[0159] Table 15 Single-plex QPCR test results of target pathogen primers
[0160]
[0161] Table 15 (continued) Single-plex QPCR test results of target pathogen primers
[0162]
[0163] Numerous experiments have shown that primers (pairs) with excellent or good amplification efficiency can function in multiplex PCR systems and effectively amplify target pathogens. Therefore, the single-pair PCR validation of the primers screened in Example 1 also passed.
[0164] Example 3 Wet-lab validation of the multiple primer system in real clinical samples
[0165] To further validate the effectiveness of the present invention's design and optimization of multiplex primer systems, the mixed system was further validated in wet-tests using the three target pathogen-specific primers described above in real clinical samples. Due to the complexity of real clinical cerebrospinal fluid samples and their generally low pathogen concentrations, the wet-test validation in this example further demonstrated the high efficiency and sensitivity of the primers designed in this invention.
[0166] (1) Test reagents and materials
[0167] Positive human cerebrospinal fluid sample known to contain the target pathogen; in-house developed nucleic acid extraction kit; reverse transcription kit (Abclonal, Cat. No. RM20428); first-round amplification kit (TaKaRa, Cat. No. RR062B); second-round amplification kit (Vazyme, Cat. No. N616-02); dsDNA quantification kit (Yeason, Cat. No. 12640ES76); ssDNA quantification kit (Yeason, Cat. No. 12645ES76); Hieff NGS DNA Selection Beads (Yeason, 12601ES75); DNBSEQ-E25RS High-Throughput Sequencing Reagent Set (FCL SE100) (MGI, Cat. No. 940-000573-00); UltraPure DNase / RNase-Free Distilled Water (Invitrogen, Cat. No. 10977015), etc.
[0168] (2) Test instruments
[0169] TL2010S medium-throughput tissue grinder (Dinghaoyuan Technology, Catalog No. 0401265); automatic nucleic acid extractor (hema, Catalog No. E96-II); PCR instrument (Longji, Catalog No. A300); Qubit 3.0 fluorescent quantitative instrument (Thermo Fisher, Catalog No. Q33216); DNBSEQ-E25RS gene sequencer (MGI, Catalog No. 900-000473-00); vortex mixer Vortex 5 (Qilin Bell, Catalog No. HD201805135); MiniL-12G large-capacity mini centrifuge (Maigao, Catalog No. MinilL-12G), etc.
[0170] (3) Primer synthesis and preparation
[0171] Based on the primers synthesized in Example 2, the second-round primers used were the Novozymes NM344 kit.
[0172] (4) Extraction of nucleic acid from samples
[0173] The samples used in the present invention are provided by the units to which this application belongs. These samples are real clinical samples confirmed to contain at least one target pathogen. An internal standard was added to the cerebrospinal fluid sample, mixed, centrifuged, and the supernatant (about 400 μL of residual liquid) was discarded. After the residual liquid was broken, lysate LE and proteinase K were added thereto. After pipetting and mixing, about 600 μL of supernatant was drawn after brief centrifugation and added to the 1 / 7 row of wells in a 96-well deep-well plate, and nucleic acid extraction was completed using a self-developed nucleic acid extraction kit. The extracted nucleic acid was then concentrated using Qubit dsDNA reagent.
[0174] (5) PCR amplification
[0175] cDNA synthesis and library preparation were performed for the aforementioned nucleic acids using a reverse transcription kit (Abclonal), a first-round amplification kit (Takara), and a second-round amplification kit (Vazyme). Product concentrations were determined using Qubit dsDNA reagent after each round of amplification.
[0176] (6) Sequencing
[0177] The libraries were prepared using nanospheres (DNBs) and sequenced using the DNBSEQ-E25RS High-Throughput Sequencing Kit. Single-end 50bp sequencing was performed. Equal weights of the libraries were mixed, and DNB preparation was performed using a 100 fmol input according to the instrument's instructions. Immediately, 20 μL of DNB termination buffer was added, and the mixture was mixed by slowly pipetting with a widened pipette tip. The single-stranded DNA concentration of the DNBs was determined using the Qubits sDNA Assay Kit. Once the concentration reached a range of 4 ng / μL to 40 ng / μL, sequencing was performed on the E25 instrument.
[0178] (7) Data analysis
[0179] After the machine is unloaded, the raw data is a FastQ file, which is analyzed using an automated bioinformatics process. The FastQ file generates results after quality control, primer matching, and Kraken 2 species identification process. The specific counting method for a target pathogen-specific Reads is: all Reads match the upstream primers and UMI sequences of the target pathogen, and identify the species in Kraken 2. Reads that are 100% matched with the upstream primers of the target pathogen and correctly report the target pathogen or target pathogen atomic level in Kraken 2 are then corrected by the number of UMI species connected to it and the number of cleanreads after quality control filtering of the data set. If the corrected number of reads of a species is >= 100 times the number of reads of the species in the blank sample, the species is considered positive in the sample. The calculation formula for the corrected number of reads is as follows:
[0180] (UMI-normalizedReads / cleanreads)*1000000
[0181] (8) Correction Reads results of target pathogens:
[0182] Qualified primers for the three target pathogens were mixed into a single system and tested using the above steps in clinical samples containing one of the target pathogens. A total of three positive samples (samples 1-3) and one blank control sample (sample 4) were tested. The test results are shown in Table 16. The positive samples only show the corrected read counts for the species known to be present. The blank control shows the corrected read counts for the three target pathogens for ease of calculation:
[0183] Table 16 Detection results of target pathogen primer mixture system in real clinical samples
[0184]
[0185] In summary, the hybrid system detected all target species contained in three positive real clinical samples, verifying the effectiveness of the primers designed and optimized by the present invention in actual clinical applications.
Claims
1. A primer design method for multiplex PCR targeted sequencing technology, characterized in that: The method comprises the following steps in sequence: S100: Determine the optimal reference genome and construct a high-quality genome library of the target pathogen; S200: constructing a phylogenetic tree for the high-quality genome library, and meeting the following requirements: the branches of the genome evolutionary tree are uniform, with neither isolated branches with significantly distant evolutionary distances nor a situation where the majority of genomes are densely distributed on the same branch; S300: traversing the optimal reference genome with a fixed step size to obtain candidate primer pairs; S400: Evaluate the candidate primer pairs, including scoring of coverage, specificity, and dimer binding; The step S100 specifically includes the following steps: S110: Obtain the best reference genome of the target pathogen; S120: Acquire several genomes of other species that were sequenced most recently, wherein the genomes of other species are acquired in descending order of assembly level; S130: Screening for other genomes of the same species with high integrity and meeting the sample quantity. If the sample quantity is insufficient, the sequencing time of step S120 is preferentially extended. The coverage evaluation method is to match the candidate primer pairs to each other genome of the same species in the high-quality genome library and calculate the primer coverage: Primer coverage = number of qualified matching genomes / total number of genomes × 100%; A qualified primer coverage match satisfies at least the following conditions simultaneously: both the upstream primer and the downstream primer can achieve a match, the matching length of the primer amplicon is not less than 95% of the total amplicon length, and the primer amplicon length is 100-300bp.
2. The primer design method for multiplex PCR targeted sequencing technology according to claim 1, characterized in that: The method for determining the optimal reference genome is as follows: screening an optimal reference genome from the RefSeq database or the GenBank database in the order of priority of Representative Genome>Reference Genome>Complete Genome>Chromosome>Scaffold>Contig; When there are multiple genomes or genome structures at the same level, the genome or genome structure with the most recent sequencing year and the highest CheckM value is selected; Among them, the sequencing year has a higher priority than the CheckM value; and if the sequencing year and CheckM value of multiple genomes or genome structures of the same level are the same, one is selected arbitrarily.
3. The primer design method for multiplex PCR targeted sequencing technology according to claim 1 or 2, characterized in that: For pathogens classified as "species" in taxonomy, the number of samples must be no less than 10; For pathogens classified as "species complex" or "genus" in taxonomy, the number of samples is 50.
4. The primer design method for multiplex PCR targeted sequencing technology according to claim 3, characterized in that: CheckM2 was used to perform integrity testing and select genomes with integrity greater than or equal to 95%.
5. The primer design method for multiplex PCR targeted sequencing technology according to claim 1, characterized in that: In step S300, the whole genome is sliced in a shingled manner with a sliding window length of 300 bp and a sliding step length of 20 bp, and primers are designed on each fragment using primer3 software.
6. The primer design method for multiplex PCR targeted sequencing technology according to any one of claims 1-2, 4-5, characterized in that: The specific evaluation method comprises the following steps: S421: aligning the candidate primer pairs in pairs in the NCBI nt library against all genomes in the library; S422: After alignment, a primer insert is obtained. The primer insert refers to a nucleic acid sequence defined and amplified by a pair of complementary oligonucleotide primers during PCR. The starting end of the primer insert is defined by the 5' end of the forward primer, and the ending end is defined by the 3' end of the reverse primer. S423: Inputting the primer inserts into Kraken 2 software for verification, thereby screening primer pairs with qualified specificity; S424: When the species of the genome to which the primer insert fragment belongs is consistent with the primer target species, the species discrimination result of Kraken 2 must be that species; or when the species of the genome to which the primer insert fragment belongs is inconsistent with the primer target species, the species discrimination result of Kraken 2 must be consistent with the genome species.
7. The primer design method for multiplex PCR targeted sequencing technology according to claim 6, characterized in that: The dimer binding score is: When the Score of any primer in the candidate primer pair is not less than 10 3 When , the primer pair corresponding to the primer is screened out; Where GC-count represents the number of GC bases in the dimer binding region, L represents the base length of the dimer binding region, d1 represents the distance from the start point of the dimer binding region to the 3' end of the nearest primer, and d2 represents the distance from the end point of the dimer binding region to the 3' end of the nearest primer.
8. The primer design method for multiplex PCR targeted sequencing technology according to claims 1-2, 4-5, and 7, characterized in that: The method is used for multiplex PCR targeted pathogen detection.
9. The primer design method for multiplex PCR targeted sequencing technology according to claim 8, characterized in that: The pathogens include bacteria, fungi and viruses.
Citation Information
Patent Citations
Design method of microbial multiplex PCR (Polymerase Chain Reaction) targeted amplification primer pool
CN115678967A
Screening method of pathogenic microorganism metagenome phylogenetic tree
CN117316263A
Probe design method, electronic equipment and computer readable storage medium
CN118782147A
Cited By
Real-time fluorescent quantitative PCR (polymerase chain reaction) primer and probe design method based on panogenomics
CN121674531A