A method, system and storage medium for designing and evaluating whole-virus primers
By designing primers based on representative genomes or conserved regions in the viral genome sequence library and conducting specificity verification, the problems of universality and accuracy in viral primer design were solved, and highly specific identification of viral subtypes and low-risk application were achieved.
Patent Information
- Application Number
- CN202510436102.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-04-09
AI Technical Summary
Existing viral primer design methods lack universality and accuracy, making it difficult to effectively identify viral subtypes and posing the risk of nonspecific amplification.
Use Scheme 1 or 2 to design whole-virus primers. Determine representative genomes or conserved regions in the viral genome sequence library, design primers, and perform specificity verification to ensure that the primer verification pass rate between species and hosts meets the standards.
It achieves high specificity and high accuracy in identifying different viral subtypes, reduces the risk of primer application, and is suitable for a variety of viruses and host environments.
Smart Images

Figure CN119932234B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of primer design, and in particular to a method, system and storage medium for designing and evaluating whole-virus primers. Background Art
[0002] Viruses, as common pathogens, have a significant impact on human health and the agricultural production of plants and animals. Timely and accurate identification of viral species and subtypes facilitates the development of precise prevention and control strategies. Traditional primers for single-virus or multi-virus identification often fail to effectively meet the needs of comprehensive viral spectrum monitoring. Recently, targeted next-generation sequencing (tNGS) has enabled the simultaneous detection of more than 40 common clinical pathogens ([Almas, S.; Carpenter, RE; Singh, A.; Rowan, C.; Tamrakar, VK; Sharma, R. Deciphering Microbiota of Acute Upper Respiratory Infections: A Comparative Analysis of PCR and mNGSMethods for Lower Respiratory Trafficking Potential. Adv. Respir. Med. 2023, 91, 49-65.]). However, limitations exist, such as false positives due to contamination caused by poor primer specificity and insufficient viral subtype identification. High specificity and efficiency are key characteristics of high-quality primers and are the goals of primer design.
[0003] Typically, when designing primers for pathogenic viruses, software is used to first calculate the conserved regions of the virus, then primers are designed for these regions, and finally, the primer specificity index is calculated. Taking the mainstream Primer-BLAST software as an example, when designing primers for a single template, its length limit is 50kb. This software can also design primers for a group of sequences. The specific operation process is as follows: first, the longest sequence is blasted against other sequences to identify conserved regions. Next, the conserved regions on the longest sequence are extracted and used as representative template sequences. Primer3 is then called for primer design. In addition, primer specificity testing can specify genome databases (such as RefSeq representative genomes) or specific species genomes (such as Home sapiens) to calculate nonspecific amplification efficiency.
[0004] Existing viral primer design methods can basically meet the needs of single virus or multi-virus detection, but they also have certain limitations. The details are as follows:
[0005] (1) The design of primers for viral subtype identification is not universal: The subtype classification of each virus is unique and requires personalized primer design. Currently, there is no method or system to design primers for different subtypes of a virus.
[0006] (2) Low accuracy of primer design: There is a lack of comprehensive consideration of the variation regions between different strains of the same species, so when the primers are designed in these variation regions, amplification failure is very likely to occur for other strains; because the mutation and evolution of viruses are very rapid, primers are currently designed only for a specific virus, and the reference information is incomplete, resulting in a high risk of application of the calculated primers. The specificity test of the designed primers is not optimal, which limits the application of primers across species. The specificity in the host is unclear, resulting in an unknown risk of nonspecific amplification in the application, which in turn makes the application risk high.
[0007] Therefore, there is an urgent need to develop a method and system for the design and evaluation of whole-virus primers. Summary of the Invention
[0008] In view of the above problems, the present invention provides a method, system and storage medium for designing and evaluating whole-virus primers.
[0009] The technical solutions adopted by the present invention to solve the technical problems are as follows:
[0010] In a first aspect, the present invention provides a method for designing and evaluating whole-virus primers, comprising:
[0011] obtaining a viral genome sequence library of the virus;
[0012] Use Scheme 1 or 2 to design usable primers for species identification, perform interspecies verification and host verification on the usable primers for species identification, and evaluate each pair of usable primers for species identification;
[0013] Use Scheme 1 or Scheme 2 to design usable primers for subtype identification, perform non-target virus subtype verification, interspecies verification, and host verification on the usable primers for subtype identification, and evaluate each pair of usable primers for subtype identification.
[0014] The first scheme is: determining a representative genome of the virus in a viral genome sequence library, designing primers for the representative genome, using non-representative genomes in the viral genome sequence library to perform specificity verification on the primers for the representative genome, and selecting the primers for the representative genome that meet the verification pass rate as usable primers;
[0015] The second scheme is: perform multiple sequence alignment on the genome sequences in the viral genome sequence library to identify the conserved regions of the virus, design primers for the conserved regions, use the viral genome sequence library to perform specificity verification on the primers in the conserved regions, and use the primers in the conserved regions with a verification pass rate that meets the standard as available primers.
[0016] In a preferred embodiment, the viral genome sequence library is a processed viral genome sequence library, and the viral genome sequence library of the virus is obtained specifically by: obtaining the virus, collecting the subtype information, genome sequence and genome sequence metadata of the virus genome to form an original viral genome sequence library, and performing sequence deduplication and filtering on the original viral genome sequence library to obtain a processed viral genome sequence library.
[0017] In a preferred embodiment, if the virus is a single-genome virus, Scheme 1 is used to design available primers for species identification and available primers for subtype identification; if the virus is a non-single-genome virus, Scheme 1 or Scheme 2 is used to design available primers for species identification based on the variability of the non-single-genome virus and / or the number of genomes of the non-single-genome virus, and Scheme 1 or Scheme 2 is used to design available primers for subtype identification based on the variability of the non-single-genome virus and / or the number of genomes of the non-single-genome virus.
[0018] In a preferred embodiment, the specificity verification specifically comprises: performing specificity analysis on the primers representing the genome using the verification genome, determining whether the primers representing the genome meet the specificity standard to obtain a specificity verification result, and calculating the specificity verification pass rate of the primers representing the genome based on the specificity verification result;
[0019] In the first scheme, the verification genome is a non-representative genome in the viral genome sequence library. If the virus does not have a non-representative genome, the primers representing the genome are considered to have passed specific verification, and the verification pass rate is 100%. If the virus has a non-representative genome, the non-representative genome in the viral genome sequence library is used as the verification genome;
[0020] In the second scheme, the verification genome is the genome of the viral genome sequence library.
[0021] In a preferred embodiment, the specificity criterion is to meet all of the following conditions:
[0022] Condition 1: The primers can amplify the genome and the amplification position is unique;
[0023] Condition 2: The gene amplified by the primers on the verification genome is consistent with the gene amplified by the primers on the representative genome. If the verification genome is a segment genome, the amplification product of the primers on the verification genome and the amplification product of the primers on the verification genome are located in the same segment;
[0024] Condition 3: Verify that the length of the amplified product on the genome and the length variation of the amplified product of the primer on the representative genome are within the set length variation threshold;
[0025] Condition 4: Verify that the sequence similarity between the amplified product on the genome and the amplified product of the primer on the representative genome reaches the set sequence similarity threshold condition.
[0026] In a preferred embodiment, the inter-species and host validation of the primers available for species identification comprises:
[0027] Amplification of the available primers for species identification in the sequences of the representative viral sequence database, and statistical amplification to obtain the viral species verification results of the available primers for species identification, wherein the representative viral sequence database includes representative sequences of multiple viral species,
[0028] Primers available for species identification are amplified on all host genomes in a host genome database, and the amplification conditions are statistically analyzed to obtain host verification results of the primers available for species identification, wherein the host genome database includes genomes of several hosts of the virus.
[0029] In a preferred embodiment, the non-target virus subtype verification, virus species verification and host verification of the available primers for subtype identification specifically include:
[0030] The available primers for subtype identification are amplified on the genome of the non-target subtype, and the amplification results are statistically analyzed to obtain the non-target virus subtype verification results;
[0031] The available primers for subtype identification are amplified on all representative sequences in a representative viral sequence database, and the amplification results are statistically analyzed to obtain viral species validation results of the available primers for subtype identification, wherein the representative viral sequence database includes representative sequences of multiple viral species;
[0032] The primers available for subtype identification are amplified on all host genomes in the host genome database, and the amplification conditions are statistically analyzed to obtain host verification results of the primers available for subtype identification. The host genome database includes genomes of several hosts of the virus.
[0033] In a preferred embodiment, the evaluation of each pair of available primers for species identification is specifically as follows: for each pair of available primers for species identification, a quality rating is performed based on its verification pass rate, its viral species verification result, and its host verification result;
[0034] The evaluation of each pair of available primers for subtype identification is specifically as follows: for each pair of available primers for subtype identification, quality rating is performed based on its verification pass rate, its non-target virus subtype verification result, its virus species verification result and its host verification result.
[0035] In a second aspect, the present invention provides a system for designing and evaluating whole-virus primers, comprising:
[0036] The first design evaluation module is used to design usable primers for species identification using Scheme 1 or Scheme 2, perform interspecies verification and host verification on the usable primers for species identification, and evaluate each pair of usable primers for species identification;
[0037] The second design evaluation module is used to design primers for subtype identification using Scheme 1 or Scheme 2, perform non-target virus subtype verification, interspecies verification, and host verification on the primers for subtype identification, and evaluate each pair of primers for subtype identification;
[0038] The first scheme is: determining a representative genome of the virus in a viral genome sequence library, designing primers for the representative genome, using non-representative genomes in the viral genome sequence library to perform specificity verification on the primers for the representative genome, and selecting the primers for the representative genome that meet the verification pass rate as usable primers;
[0039] The second scheme is: perform multiple sequence alignment on the genome sequences in the viral genome sequence library to identify the conserved regions of the virus, design primers for the conserved regions, use the viral genome sequence library to perform specificity verification on the primers in the conserved regions, and use the primers in the conserved regions with a verification pass rate that meets the standard as available primers.
[0040] In a third aspect, the present invention provides a storage medium storing a computer program, which, when executed by a processor, implements the steps of the method for designing and evaluating a whole-virus primer described in the first aspect.
[0041] The present invention is a method and system for the joint development of primers for viral species identification and viral subtype identification. This method obtains usable primers through primer design and specificity verification of the designed primers, including the design of corresponding primers for different viral subtypes. The method comprehensively verifies usable primers for species identification and subtype identification, providing analysis and evaluation based on intra-species, inter-species, and (inter-subtype) host perspectives. This design makes the method applicable to viral primer design for diverse application scenarios and with varying genomic characteristics. It offers universal applicability and high accuracy for viral subtype identification primer design. On a full viral scale, it ensures high specificity of candidate viral primers, improves primer accuracy, and reduces the risk of primer application. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 A flowchart of a method for designing and evaluating whole-virus primers;
[0043] Figure 2 Flowchart for the design and evaluation of useful primers for viral species identification;
[0044] Figure 3 Flowchart for designing primers available for viral species identification using Scheme 1;
[0045] Figure 4 Flow chart for validation of usable primers for species identification using the protocol;
[0046] Figure 5 Flowchart for designing primers that can be used for viral species identification using Scheme 2;
[0047] Figure 6 Flowchart for the design and evaluation of useful primers for viral subtype identification;
[0048] Figure 7 A framework diagram of a system for designing and evaluating whole-virus primers. DETAILED DESCRIPTION
[0049] In order to more clearly understand the above-mentioned objects, features and advantages of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0050] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the scope of protection of the present invention is not limited to the specific embodiments disclosed below.
[0051] It should be noted that the descriptions of "first", "second", etc. in this application are for descriptive purposes only and should not be understood as indicating or implying their relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" or "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between the various embodiments can be combined with each other, but this must be based on the fact that they can be implemented by ordinary technicians in this field. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by this application.
[0052] The existing methods and systems for designing viral primers not only lack universal applicability, but also have low accuracy and high application risks. To this end, this application provides a method for designing and evaluating all-viral primers, such as Figure 1 , the method comprising:
[0053] obtaining a viral genome sequence library of the virus;
[0054] Design usable primers for species identification and / or design usable primers for subtype identification.
[0055] The method of designing the available primers for species identification comprises the following steps: designing the available primers for species identification using scheme 1 or scheme 2, performing interspecies verification and host verification on the available primers for species identification, and evaluating each pair of available primers for species identification;
[0056] The available primers for subtype identification are designed as follows: using scheme one or scheme two to design available primers for subtype identification, performing non-target virus subtype verification, inter-species verification and host verification on the available primers for subtype identification, and evaluating each pair of available primers for subtype identification.
[0057] The first scheme is: determining the representative genome of the virus in the viral genome sequence library, designing primers for the representative genome, using the non-representative genome in the viral genome sequence library to perform specificity verification on the primers of the representative genome, and the primers of the representative genome that meet the specificity verification pass rate are used as available primers.
[0058] The second scheme is: perform multiple sequence alignment on the genome sequences in the viral genome sequence library to identify the conserved regions of the virus, design primers for the conserved regions, use the viral genome sequence library to perform specificity verification on the primers in the conserved regions, and use the primers in the conserved regions with a specificity verification pass rate that meets the standard as available primers.
[0059] Here, the order of the primers available for species identification and the primers available for subtype identification is not limited.
[0060] In this embodiment, the viral genome sequence library is a processed viral genome sequence library, and the viral genome sequence library of the virus is obtained by: obtaining the virus, collecting the subtype information, genome sequence and genome sequence metadata of the virus genome to form an original viral genome sequence library, and performing sequence deduplication and filtering on the original viral genome sequence library to obtain a processed viral genome sequence library.
[0061] The method of this embodiment is divided into two parts: design and evaluation of primers for viral species identification and design and evaluation of primers for viral subtype identification. The order of the two parts is not limited.
[0062] Part I: Design and evaluation of primers for viral species identification. Figure 2 , including the following steps S11~S15.
[0063] S11. Obtain a virus (a virus for which primers are to be designed), collect the genome sequence and genome sequence metadata of the virus for which primers are to be designed to form an original virus genome sequence library 1, and screen the genome sequences in the original virus genome sequence library 1 to obtain a processed virus genome sequence library 1.
[0064] The screening of the genome sequence is to perform sequence deduplication and filtering on the original viral genome sequence library.
[0065] Specifically, all the genome sequences and metadata of the virus are collected from public databases to form the original viral genome sequence library 1, and then sequence deduplication and low-quality sequence filtering are performed, so that the original viral genome sequence library 1 can retain valid information while streamlining the amount of sequence data, thereby improving the efficiency of primer design.
[0066] As an embodiment, the sequence deduplication is specifically performed by using cd-hit software to retain only one 100% identical genomic sequence. The sequence filtering is performed to remove short fragment sequences, and the specific length threshold for filtering can be designed according to actual conditions.
[0067] S12. Design and verify primers for virus species identification to obtain usable primers for species identification.
[0068] S12.1. Determine whether S12.2 or S12.3 is used for virus species identification primer design based on whether the virus is a single-genome virus. This determination is also based on the variability of non-single-genome viruses and / or the number of genomes of non-single-genome viruses. Variability, as used herein, can be understood as the mutation rate.
[0069] The present invention provides two schemes for designing primers for viral species identification. For single-genome viruses, scheme one (S12.2) is directly selected to design species identification primers. Since it does not involve multiple sequence alignment, the time cost and computational cost are lower than those of scheme two. For non-single-genome viruses, scheme two (S12.3) can be directly selected to design primers. Conserved regions are first extracted and then species identification primers are designed. This can more quickly obtain primers with high species specificity, but it is not limited to non-single-genome viruses that must choose scheme two.
[0070] Typically, a variability threshold and a genome number threshold can be set to determine whether option 1 or option 2 is selected for a non-single genome virus based on the variability threshold and / or genome number threshold of the non-single genome virus.
[0071] In one embodiment, when the variability of the non-monogenome virus exceeds the variability threshold and the number of genomes exceeds the genome number threshold, option 2 is selected, otherwise option 1 is selected; in one embodiment, when at least one of the variability of the non-monogenome virus exceeds the variability threshold and the number of genomes exceeds the genome number threshold is true, option 2 is selected, otherwise option 1 is selected; in one embodiment, variability or genome number is not considered, for example, variability is not considered, when the number of genomes of the non-monogenome virus exceeds the genome number threshold, option 2 is selected, otherwise option 1 is selected; for example, genome number is not considered, when the variability of the non-monogenome virus exceeds the variability threshold, option 2 is selected, otherwise option 1 is selected. It will be understood that the above embodiments are merely examples and not limitations.
[0072] As a specific example, for viruses whose genome length is less than 100 kb for which primers are to be designed, the genome number threshold is set to a number less than or equal to 50. If the genome length of the virus for which primers are to be designed exceeds 100 kb, the genome number threshold is set to a number less than or equal to 10.
[0073] S12.2. Determine the representative genome in the processed viral genome sequence library 1, design species identification primers (species identification primers) for the representative genome, and obtain the species identification primer set for the representative genome; use other genomes in the processed viral genome sequence library 1 (other genomes are non-representative genomes) to verify the species identification primers in the species identification primer set for the representative genome, screen the species identification primers for the representative genome that meet the verification pass rate as the available species identification primers, obtain the available species identification primer set, and proceed to S13.
[0074] The process of S12.2 can be found in Figure 3 .
[0075] Determine the representative genome: Select a genome from the processed viral genome sequence library 1 as the representative genome. The selection principle can be targeted strain screening based on the scientific research question, or it can prioritize RefSeq (Reference Sequence) genomes. In addition to the representative genome, the other genomes in the processed viral genome sequence library 1 are called neighbor genomes, also known as non-representative genomes, and are used to verify primers and serve as validation genomes for primer verification.
[0076] Use Primer3 software to design species identification primers for the representative genome: If all sequences in the representative genome are within 10 kb in length, primers can be designed directly using Primer3 software. However, when the representative genome is longer than 10 kb or the representative genome has contigs (or segments) longer than 10 kb, due to the length limit of Primer3 software, the long sequence (sequences that are long for Primer3 software) needs to be cut. When cutting, use a 10 kb window and a 9.5 kb step size, or use a 500 bp overlapping window mode (the product length is set between 100 bp and 300 bp, and the overlapping window is set to 500 bp to avoid primer loss due to sequence cutting). Use Primer3 software to design primers for the cut results (sub-fragments), and then merge the primer results. Deduplication must be performed when merging the primer results.
[0077] Primers that uniquely amplify positions are retained to obtain the representative genome species identification primer set. When the viral genome sequence (corresponding to the representative genome at this time) is a complete sequence, Primer3 will consider the uniqueness of the amplification position when designing primers. However, when the viral genome is composed of multiple segments, Primer3 software does not consider the specificity of primers for different segments. Therefore, it is necessary to screen for primers that uniquely amplify positions across the entire genome before they can be included in the representative genome species identification primer set.
[0078] As an example, the main parameters of primer 3 design are: primer length 18bp-23bp, product length 100bp-300bp, RCR annealing temperature 59°C, and GC content range 30%-70%.
[0079] The validation results of the species identification primers were obtained by using the neighbors genome to determine whether the virus has a neighbors genome. If the virus does not have a neighbors genome, the species identification primer set representing the genome can be directly used as the available primer set for species identification, that is, the primers of the species identification primer set representing the genome are directly considered to have passed the specificity validation, and the validation pass rate is 100%; if the virus has a neighbors genome, the species identification primer set representing the genome will be specifically tested on each neighbors genome ( Figure 3 In the , "in-house script" means "internal script" ).
[0080] The flow chart of this step is as follows Figure 4 As shown in , each neighbor genome is a validation genome. Figure 4 As shown, MEFprimer software is used to analyze the specificity of the primers to be verified on the target genome (i.e., the neighbors genome at this time), and then the specificity is determined to determine whether it meets the specificity standard to obtain the specificity verification result, and the specificity verification pass rate is calculated based on the specificity verification result. Determine whether the specificity meets the following conditions: (1) The species identification primer can have an amplification effect on the verification genome and the amplification position is unique (if there is no amplification, determine whether the genome is a segment genome and whether it has the same segment as the primer template. The purpose of this operation is to finely mark the reason for the failure of the verification); (2) The amplified gene of the species identification primer on the verification genome is consistent with the amplified gene of the species identification primer on the representative genome. If the verification genome is a segment genome, it is also required that the amplified product of the primer on the verification genome and the amplified product of the species identification primer on the verification genome are located in the same region. segment; (3) the length of the amplified product on the verification genome and the length variation of the amplified product of the species identification primer on the representative genome are within the set length variation threshold range (such as 10%); (4) the sequence similarity of the amplified product on the verification genome and the amplified product of the species identification primer on the representative genome reaches the set sequence similarity threshold condition (such as the blast comparison result shows identity ≥ 95%, coverage ≥ 95%, identity is the ratio of the consistent base sites in the two sequences to the total number of bases; coverage is the ratio of the length of the comparison region to the total sequence length). If all the above conditions are met, it is considered to meet the specificity standard, and the specificity verification result of the primer pair on a verification genome is verified. Specificity analysis and judgment are performed one by one to obtain a verified representative genome species identification primer set.
[0081] The validated representative genome species identification primer sets are screened for viral species identification primers to obtain a usable primer set for species identification. The available primers are screened by setting a pass rate threshold. Specifically: After validation with neighbor genomes, the pass rate of each representative genome species identification primer pair relative to all validated genomes is calculated using the optional "rounding off" principle. This is called the validation pass rate. Species identification primers for representative genomes that meet the pass rate threshold (e.g., 0.5) are then selected as usable primers for species identification, forming the usable primer set for species identification.
[0082] S12.3, see Figure 5 , perform multiple sequence alignment on the genome sequences in the processed viral genome sequence library one, the number of genome sequences in the processed viral genome sequence library one needs to be greater than or equal to 2, identify the conserved regions of the virus based on the multiple sequence alignment results, design species identification primers for the conserved regions, and obtain a species identification primer set for the conserved regions, use the genome in the processed viral genome sequence library one to verify the primers in the species identification primer set for the conserved regions, screen the species identification primers for the conserved regions with a verification pass rate that meets the standard as available species identification primers, obtain a available species identification primer set, and proceed to S13.
[0083] Multiple sequence alignment: Clustal OMEGA software (default parameters) can be used to perform multiple sequence alignment on partial or complete genome sequences from the processed viral genome sequence library to obtain the sequence. If the genome is a segment genome, the multiple sequence alignment is performed based on the "same segment" or "same gene" in different genomes.
[0084] Conserved region identification: Gblocks software was used, and its important parameter settings were as follows:
[0085] -b1 The minimum support number of single bases in the conserved region is set to the smallest integer that is not less than 50% of the genome number;
[0086] -b2 The minimum support number of single bases at the boundary of the conserved region is set to the minimum integer that is not less than 85% of the genome number;
[0087] -b3 The maximum number of consecutive non-conservative bases is set to 8bp;
[0088] -b4 The minimum length of the conserved region is set to 200 bp;
[0089] -b5 Whether to allow gap positions (None, With Half, All). Set to a to allow.
[0090] Extract the conserved region sequence as the template sequence: traverse each base position in the conserved region, and retain the base with the highest support number at each position (if two bases have the same support number, the bases are sorted alphabetically first). If the base with the highest support number is "-", skip it.
[0091] Based on the extracted conserved region sequence, primer 3 software was used to design primers for the conserved region. The specific process can refer to the design of primers for representative genomic species identification, except that the input file was changed from the representative genomic sequence to the conserved region sequence.
[0092] The species identification primers were verified using the genomes of the processed viral genome sequence library one to obtain verification results: the species identification primer set in the conserved region was specifically tested on each genome of the processed viral genome sequence library one (i.e., specificity verification).
[0093] As an embodiment, each genome in the processed viral genome sequence library 1 is a verification genome. As another embodiment, all genomes in the processed viral genome sequence library 1 except the genomes with the template sequence are used as verification genomes and can be considered as neighbors genomes. MEFprimer software was used to analyze the specificity of the species identification primers on the target genome (verification genome), and then the results were tested to see if they met the following conditions: (1) the species identification primers had an amplification effect on the verification genome and the amplification position was unique; (2) the gene amplified by the species identification primers on the verification genome was consistent with the gene amplified by the species identification primers in the conserved region. If the verification genome was a segment genome, the amplification product of the primers on the verification genome and the amplification product of the species identification primers on the verification genome were also required to be located in the same segment; (3) the length of the amplification product on the verification genome and the length of the amplification product of the species identification primers in the conserved region varied within a set threshold range (e.g., 10%); (4) the sequence similarity between the amplification product on the verification genome and the amplification product of the species identification primers in the conserved region reached a set threshold condition (e.g., blast comparison results showed identity ≥ 95%, coverage ≥ 95%, identity was the ratio of the number of consistent base sites in the two sequences to the total number of bases; coverage was the ratio of the length of the comparison region to the total sequence length). If all the above conditions are met, the test is considered to have passed, and a species identification primer set of the verified conserved region is obtained.
[0094] The species identification primer sets of the verified conserved regions are screened to obtain the usable primer sets for species identification: After the genome verification is completed, according to the principle of "rounding off and leaving even numbers", the pass rate of each pair of species identification primers in the conserved regions relative to all verified genomes is calculated (statistically), which is called the verification pass rate. Then, the species identification primers in the conserved regions that meet the pass rate threshold (for example, 0.5) are screened as the usable primers for species identification, forming the usable primer set for species identification.
[0095] S13. The available primer set for species identification is amplified on all representative sequences in the virus representative sequence database to obtain the amplification status of the available primers for species identification in the sequences in the virus representative sequence database, and the amplification status (the number of amplified targets and species information of each pair of available primers for species identification on the genomes of other species) is counted to obtain the virus species verification results. The virus representative sequence database includes representative sequences of multiple virus species.
[0096] S13.1. Construct a representative viral sequence database.
[0097] For each virus species, all its sequences are collected and then de-redundanted, and its representative sequences are retained. The representative sequences of each species constitute the virus representative sequence database.
[0098] The process of building a representative viral sequence database includes three key steps:
[0099] ① Metadata cleaning and standardization: Standardize metadata such as host, country, time, and separation source. Delete duplicate metadata sequences.
[0100] ② Sequence clustering and quality control: MMseq2 was used to cluster sequences by species, and sequences that were 100% covered by other fragments were deleted. BlastN was used for pairwise sequence alignment to remove sequences that could not be properly clustered.
[0101] ③ Representative sequence screening: retain the sequences with the most complete metadata and sequence information for each branch. These sequences are used as representative sequences. The representative sequences of each viral species constitute the viral representative sequence database.
[0102] S13.2. Perform interspecies verification on the primers available for species identification to obtain interspecies verification results of the virus.
[0103] Use MEFPrimer software with default parameters, input the available primer set for species identification and the representative viral sequence database, and MEFPrimer software will output the amplification status of the available primers for species identification in the sequences of the representative viral sequence database, count the number of amplified targets and species information of each pair of available primers for species identification on the genomes of other species, and obtain the viral species verification result, which is called viral species verification result 1, or the viral species verification result of the available primers for species identification, for subsequent primer evaluation.
[0104] S14. Amplify the available primer set for species identification on all host genomes in the host genome database, count the amplification results, and obtain host verification results. The host genome database includes the genomes of several hosts of the virus (corresponding to the available primers for species identification).
[0105] Specifically, S14 includes:
[0106] S14.1. Construct a host genome database of the virus.
[0107] For each virus, we search for its common hosts using DNA sequence databases (such as NCBI) or literature. For common hosts, we can define what constitutes commonness. For example, if a particular host accounts for more than 10% of all known hosts of a virus, we define it as common. We download the genome sequences of common hosts and create a host genome database. Each virus has a host genome database.
[0108] S14.2. Perform host validation on the primers available for species identification.
[0109] MEFPrimer software was used with default parameters. The software inputted the primer set for species identification and the corresponding host genome database. The software then outputted the amplification results of the primers for species identification in the host genome. The number of amplified targets and host species information for each pair of primers for species identification were counted to generate host validation results for subsequent primer evaluation.
[0110] It can be understood that the order of S13 and S14 is not limited.
[0111] S15. Perform a quality rating on each pair of available primers for species identification in the set of available primers for species identification. Specifically, perform a quality rating on each pair of available primers for species identification based on its (the pair of available primers for species identification) verification genome pass rate (the verification pass rate), its viral species verification result (viral species verification result one), and its host verification result (host verification result one).
[0112] As an example and not a limitation, let factor A represent the pass rate of the verified genome (if there is no verified genome A, it is set to 100%), factor B represents the number of other viral species that can be amplified, and factor C represents the number of host genomes that can be amplified. The primer evaluation reference scheme can be referred to Table 1.
[0113] Table 1
[0114]
[0115] Level L1 primers are classified as having no validated genome or a 100% validation genome pass rate, and have no amplification effect on genome sequences of other viral species or the host genome. These primers are considered the most specific for species identification.
[0116] Level 2: While the genome validation pass rate is less than 100%, it exceeds 50%, and there is no amplification effect on the genome sequences of other viral species or the host genome. This type of primer set is considered highly specific for species identification.
[0117] Level L3-L4: A genome validation pass rate of at least 50% is achieved, and amplification is observed on no more than two non-native viral genomes. Primers are categorized as L3 or L4 based on whether or not they amplify the host genome. When using primers at this level, consider the application scenario and the coexistence of contaminating viruses and hosts in the sample being tested.
[0118] Levels L5-L6: A genome validation pass rate of at least 50% and amplification activity on at least two viral genomes from non-native species. Levels L5 and L6 are categorized based on whether or not amplification activity is observed on the host genome. When using these level primers, caution should be exercised based on the application scenario and requirements.
[0119] The second part is the design and evaluation scheme of primers available for virus subtype identification. The process framework is shown in the figure below. Figure 6 The design and evaluation of primers for viral subtype identification includes the following steps S21 to S25.
[0120] S21. Obtain a virus (the virus for which primers are to be designed), collect the genome sequence and genome sequence metadata of the virus for which primers are to be designed, and the subtype information of the virus genome to form an original virus genome sequence library 2, and screen the genome sequences in the original virus genome sequence library 2 to obtain a processed virus genome sequence library 2.
[0121] The genome sequence screening is to perform sequence deduplication and filtering on the original viral genome sequence library. The specific process can be found in S11.
[0122] It can be understood that the design of available primers for species identification and the design of available primers for subtype identification share the same viral genome sequence library. Here, S11 and S21 can be the same step. In some embodiments, only S21 is performed and S11 is not performed, that is, the processed viral genome sequence library 2 is used in S12 to S14.
[0123] S22. Extract the genome set of the target subtype, design and verify primers for subtype identification to obtain usable primers for subtype identification.
[0124] If we know which genes or fragments determine different subtypes through literature research or prior knowledge, we can first extract the variant genes or fragments of each genome and then design primers for subtype identification.
[0125] S22.1. Determine whether to use S22.2 or S22.3 for viral subtype identification primer design based on whether the virus is a single-genome virus. Specifically, the determination is also based on the variability (mutation rate) of non-monogeneous viruses and / or the number of genomes of non-monogeneous viruses.
[0126] For any virus, the design of primers for species identification and subtype identification can adopt the same scheme, or they can be designed with their own judgment criteria. In this embodiment, to simplify the steps, if Scheme 1 is adopted in S12, Scheme 1 is also adopted in S22, and if Scheme 2 is adopted in S12, Scheme 2 is also adopted in S22.
[0127] The present invention provides two schemes for designing primers for viral subtype identification. For single-genome viruses, scheme one (S22.2) is directly selected to design subtype identification primers. Since it does not involve multiple sequence alignment, the time cost and computational cost are lower than those of scheme two (S22.3). For non-single-genome viruses, S22.3 can be directly selected to design primers. Conserved regions are first extracted and then subtype identification primers are designed. This can obtain highly specific primers more quickly, but it is not limited to non-single-genome viruses that must choose S22.3.
[0128] Typically, a variability threshold and a genome number threshold can be set, and whether S22.2 or S22.3 is selected for a non-single genome virus is determined based on the variability threshold and / or the genome number threshold.
[0129] In one embodiment, when the variability of the non-single genome virus exceeds the variability threshold and the number of genomes exceeds the genome number threshold, S22.3 is selected, otherwise S22.2 is selected; in one embodiment, when the variability of the non-single genome virus exceeds the variability threshold and the number of genomes exceeds the genome number threshold, at least one of the above conditions is met, S22.3 is selected, otherwise S22.2 is selected; in one embodiment, variability or genome number is not considered, for example, variability is not considered, when the number of genomes of the non-single genome virus exceeds the genome number threshold, S22.3 is selected, otherwise S22.2 is selected; for example, when the number of genomes is not considered, when the variability of the non-single genome virus exceeds the variability threshold, S22.3 is selected, otherwise S22.2 is selected. It can be understood that the above embodiments are only examples and not limitations.
[0130] S22.2. Determine the representative genome in the processed viral genome sequence library 2, design primers for subtype identification of the representative genome, and obtain a subtype identification primer set for the representative genome; use other genomes (non-representative genomes) in the processed viral genome sequence library 2 to verify the subtype identification primers in the subtype identification primer set of the representative genome, screen the subtype identification primers of the representative genome that meet the verification pass rate as the available primers for subtype identification, obtain a available primer set for subtype identification, and proceed to S23.
[0131] Determine the representative genome: Select a genome from the processed viral genome sequence library 2 as the representative genome. This selection can be based on targeted strain screening based on the scientific question being studied, or by prioritizing RefSeq (Reference Sequence) genomes. In addition to the representative genome, the remaining genomes in the processed viral genome sequence library 2 are called neighbor genomes, or non-representative genomes, and are used to validate primers.
[0132] Use Primer3 software to design subtype identification primers for the representative genome: If all sequences in the representative genome are within 10 kb in length, primers can be designed directly using Primer3 software. However, when the length of the representative genome exceeds 10 kb or the representative genome has contigs (or segments) longer than 10 kb, due to the length limit of Primer3 software, long sequences (sequences that are long for Primer3 software) need to be cut. During cutting, a mode with a 500 bp overlapping window is used (the product length is set between 100 bp and 300 bp, and the overlapping window is set to 500 bp to avoid primer loss due to sequence cutting). Primers are designed using Primer3 software for the cutting results, and then the primer results are merged. Duplicate removal is required when merging the primer results.
[0133] When the viral genome sequence (corresponding to the representative genome at this time) is a complete sequence, Primer3 will consider the uniqueness of the amplification position when designing primers. However, when the viral genome consists of multiple segments, Primer3 software will not consider the specificity of the primers on different segments. Therefore, it is necessary to screen out primers with unique amplification positions on the entire genome. Only such primers can be included in the subtype identification primer set of the representative genome.
[0134] As an example, the main parameters of primer 3 design are: primer length 18bp-23bp, product length 100bp-300bp, RCR annealing temperature 59°C, and GC content range 30%-70%.
[0135] The subtype identification primers were verified using the neighbors genome and the verification results were obtained: when the virus does not have a neighbors genome, the subtype identification primer set representing the genome can be directly used as the available primer set for subtype identification, that is, the primers of the subtype identification primer set representing the genome are directly regarded as having passed the specificity verification, and the verification pass rate is 100%; if the virus has a neighbors genome, the subtype identification primer set representing the genome will be specifically tested on each neighbors genome.
[0136] Each neighbor genome is a verification genome. MEFprimer software is used to analyze the specificity of primers on the target genome (target verification genome, i.e., the neighbor genome at this time), and then test whether the results meet the following conditions: (1) The subtype identification primer can have an amplification effect on the verification genome and the amplification position is unique; (2) The amplified gene of the subtype identification primer on the verification genome is consistent with the amplified gene of the subtype identification primer on the representative genome. If the verification genome is a segment genome, it is also required that the amplified product of the primer on the verification genome and the amplified product of the subtype identification primer on the verification genome are consistent. The products are located on the same segment; (3) the length of the amplified product on the verification genome and the length of the amplified product of the subtype identification primer on the representative genome are within the set length variation threshold (such as 10%); (4) the sequence similarity of the amplified product on the verification genome and the amplified product of the subtype identification primer on the representative genome reaches the set sequence similarity threshold condition (such as the blast comparison result shows identity ≥ 95%, coverage ≥ 95%, identity is the ratio of the consistent base sites in the two sequences to the total number of bases; coverage is the ratio of the length of the comparison region to the total sequence length). If all of the above conditions are met, it is considered to meet the specificity standard, and the specificity verification result of the primer pair on a verification genome is verified. Specificity analysis is performed one by one to obtain a subtype identification primer set for the representative genome that has passed the verification.
[0137] The subtype identification primer set of the verified representative genome is screened for viral subtype identification primers to obtain a usable primer set for subtype identification: After completing the verification of the neighbors genome, according to the principle of "rounding up to the nearest integer" (optional), the pass rate of each pair of subtype identification primers of the representative genome relative to all neighbors genomes that passes the specificity standard is calculated, which is called the verification pass rate. Then, the subtype identification primers of the representative genome that meet the pass rate threshold (for example, 0.5) are screened as usable primers for subtype identification, forming the usable primer set for subtype identification.
[0138] S22.3. Perform multiple sequence alignment on the genome sequences in the processed viral genome sequence library 2, identify the conserved regions of the virus based on the multiple sequence alignment results, design subtype identification primers for the conserved regions, and obtain a subtype identification primer set for the conserved regions. Use the genomes in the processed viral genome sequence library 2 to verify the primers in the subtype identification primer set for the conserved regions, screen the subtype identification primers for the conserved regions with a verification pass rate that meets the standard as available primers for subtype identification, obtain a available primer set for subtype identification, and proceed to S23.
[0139] Multiple sequence alignment: Clustal OMEGA software (default parameters) can be used to perform multiple sequence alignment on some or all of the processed viral genome sequences in the second library to obtain the sequence information. If the genome is a segment genome, the multiple sequence alignment is performed based on the "same segment" or "same gene" in different genomes.
[0140] Conserved region identification: Gblocks software was used, and its important parameter settings were as follows:
[0141] -b1 The minimum support number of single bases in the conserved region is set to the smallest integer that is not less than 50% of the genome number;
[0142] -b2 The minimum support number of single bases at the boundary of the conserved region is set to the minimum integer that is not less than 85% of the genome number;
[0143] -b3 The maximum number of consecutive non-conservative bases is set to 8bp;
[0144] -b4 The minimum length of the conserved region is set to 200 bp;
[0145] -b5 Whether to allow gap positions (None, With Half, All). Set to a to allow.
[0146] Extract the conserved region sequence as the template sequence: traverse each base position in the conserved region, and retain the base with the highest support number at each position (if two bases have the same support number, the bases are sorted alphabetically first). If the base with the highest support number is "-", skip it.
[0147] Primer3 software was used to design primers for the conserved region sequences. The specific process can refer to the design of primers for representative genomic subtype identification, except that the input file was changed from the representative genomic sequence to the conserved region sequence.
[0148] The subtype identification primers were verified using the genomes of the processed viral genome sequence library 2 to obtain the verification results: the subtype identification primer set in the conserved region was specifically tested on each genome of the processed viral genome sequence library 2.
[0149] As an embodiment, each genome in the processed viral genome sequence library 2 is a verification genome. As another embodiment, all genomes in the processed viral genome sequence library 2 except the genomes with the template sequence are used as verification genomes and can be considered as neighbors genomes. MEFprimer software was used to analyze the specificity of the subtype identification primers on the target genome (verification genome), and then the results were tested to see if they met the following conditions: (1) the subtype identification primers had an amplification effect on the verification genome and the amplification position was unique; (2) the gene amplified by the subtype identification primers on the verification genome and the gene amplified by the subtype identification primers in the conserved region were consistent. If the verification genome was a segment genome, the amplification product of the primers on the verification genome and the amplification product of the subtype identification primers on the verification genome were also required to be located on the same segment; (3) the length of the amplification product on the verification genome and the length of the amplification product of the subtype identification primers in the conserved region varied within the set length variation threshold (e.g., 10%); (4) the sequence similarity between the amplification product on the verification genome and the amplification product of the subtype identification primers in the conserved region met the set sequence similarity threshold (e.g., blast comparison results showed identity ≥ 95%, coverage ≥ 95%, identity was the ratio of the number of consistent base sites in the two sequences to the total number of bases; coverage was the ratio of the length of the alignment region to the total sequence length). If all the above conditions are met, the test is considered to have passed, and a subtype identification primer set of the verified conserved region is obtained.
[0150] It is understandable that in different specificity verifications, the setting of the relevant threshold may be different.
[0151] The subtype identification primer sets of the validated conserved regions are screened to obtain the available primer sets for subtype identification: after the genome verification is completed, according to the principle of "rounding off and leaving even numbers", the pass rate of each pair of subtype identification primers in the conserved regions relative to all validated genomes is calculated, which is called the validation pass rate. Then, the subtype identification primers in the conserved regions that meet the pass rate threshold (for example, 0.5) are screened as the available primers for subtype identification, forming the available primer set for subtype identification.
[0152] S23. Non-target virus subtype verification: The available primer set for subtype identification is amplified on the genome of the non-target subtype, and the amplification status is counted to obtain the non-target virus subtype verification result.
[0153] Use MEFPrimer software to verify the amplification performance of the primer sets used for subtype identification on non-target subtype genomes. The number of amplified targets and subtype information for each subtype identification primer pair on the non-target subtype genomes were counted for subsequent primer evaluation. The non-target subtype genomes were obtained from processed viral genome sequence library 2.
[0154] It can be understood that, for example, if a virus has a first subtype and a second subtype, S23 is verification: verifying the amplification of the available primer set for subtype identification of the first subtype on the genome of the second subtype, and verifying the amplification of the available primer set for subtype identification of the second subtype on the genome of the first subtype.
[0155] S24. The available primer set for subtype identification is amplified on all representative sequences in the virus representative sequence database, and the amplification situation (the number of amplified targets and species information of each pair of available primers for subtype identification on the genomes of other species) is counted to obtain the virus species verification results. The virus representative sequence database includes representative sequences of multiple virus species.
[0156] S24.1. Construct a representative viral sequence database.
[0157] For each viral species, all sequences were collected and then de-redundant, retaining representative sequences. These representative sequences for each viral species constituted a representative viral sequence database. This representative viral sequence database is identical to the one in S13.1 and can be directly used without rebuilding.
[0158] The process of building a representative viral sequence database includes three key steps:
[0159] ① Metadata cleaning and standardization: Standardize metadata such as host, country, time, and separation source. Delete duplicate metadata sequences.
[0160] ② Sequence clustering and quality control: MMseq2 was used to cluster sequences by species, and sequences that were 100% covered by other fragments were deleted. BlastN was used for pairwise sequence alignment to remove sequences that could not be properly clustered.
[0161] ③ Representative sequence screening: retain the sequences with the most complete metadata and sequence information for each branch. These sequences are used as representative sequences. The representative sequences of each viral species constitute the viral representative sequence database.
[0162] S24.2. Perform inter-species validation on the primers available for subtype identification to obtain inter-species validation results.
[0163] Use MEFPrimer software with default parameters, input the available primer set for subtype identification and the representative viral sequence database (processed viral genome sequence library 2), and MEFPrimer software will output the amplification status of the subtype identification primers in the sequences of the representative viral sequence database, count the number of amplified targets and species information of each pair of subtype identification primers on the genomes of other species, and obtain the viral species verification results, which are called viral species verification results 2, or viral species verification results of available primers for subtype identification, for subsequent primer evaluation.
[0164] S25. Amplify the available primer set for subtype identification on all host genomes in the host genome database, count the amplification results, and obtain host verification results. The host genome database includes the genomes of several hosts of the virus (corresponding to the available primers for subtype identification).
[0165] Specifically, S25 includes:
[0166] S25.1. Construct a host genome database of the virus.
[0167] For each virus, its corresponding common host is searched through DNA sequence databases (such as NCBI) or literature. For common hosts, you can define what is considered common. For example, if a certain host accounts for more than 10% of all known hosts of a virus, it is defined as common. Download the genome sequences of common hosts to form a host genome database. Each virus has a host genome database. In this example, for the same virus, the host genome database for S25 and the host genome database for S14 are the same, meaning there is no need to repeatedly construct a host genome database.
[0168] S25.2. Perform host validation on the primers available for subtype identification.
[0169] Using MEFPrimer software with default parameters, input the primer set for subtype identification and the corresponding host genome database. The software will output the amplification results of the subtype identification primers in the host genome. The number of amplified targets and species information in the host for each subtype identification primer pair are counted to obtain host validation results for subsequent primer evaluation.
[0170] It can be understood that the order of S23, S24 and S25 is not limited.
[0171] S26. Perform a quality rating on each pair of available primers for subtype identification in the set of available primers for subtype identification. Specifically, for each pair of available primers for subtype identification, a quality rating is performed based on its (the pair of available primers for subspecies identification) verification genome pass rate, its non-target virus subtype verification result, its virus species verification result (virus species verification result 2) and its host verification result (host verification result 2).
[0172] As an example, not a limitation, let factor A represent the pass rate of the validated genome (if no validated genome A is set to 100%), factor B represents the number of other viral species that can be amplified, factor C represents the number of host genomes that can be amplified, and factor D represents the number of non-target subtypes that can be amplified. A reference scheme for evaluating subtype identification primers can be found in Table 2. L1 represents level 1, and similarly, L2, L3, and so on.
[0173] Table 2
[0174]
[0175] It should be understood that the above method does not represent or imply that all steps must be performed in this order. Ordinary technicians in this field can transform or change the execution order of the above steps based on the present invention. Some embodiments of the above method are illustrated below.
[0176] See Figure 7 This embodiment also provides a system for designing and evaluating whole-virus primers, the system comprising:
[0177] An acquisition module, used to obtain a viral genome sequence library of the virus;
[0178] The first design evaluation module is used to design usable primers for species identification using Scheme 1 or Scheme 2, perform interspecies verification and host verification on the usable primers for species identification, and evaluate each pair of usable primers for species identification;
[0179] The second design evaluation module is used to design primers for subtype identification using Scheme 1 or Scheme 2, perform non-target virus subtype verification, interspecies verification, and host verification on the primers for subtype identification, and evaluate each pair of primers for subtype identification;
[0180] The first scheme is: determining a representative genome of the virus in a viral genome sequence library, designing primers for the representative genome, using non-representative genomes in the viral genome sequence library to perform specificity verification on the primers for the representative genome, and selecting the primers for the representative genome that meet the verification pass rate as usable primers;
[0181] The second scheme is: perform multiple sequence alignment on the genome sequences in the viral genome sequence library to identify the conserved regions of the virus, design primers for the conserved regions, use the viral genome sequence library to perform specificity verification on the primers in the conserved regions, and use the primers in the conserved regions with a verification pass rate that meets the standard as available primers.
[0182] The first design evaluation module includes: a first design unit, a species identification verification unit, and a first evaluation unit. The first design unit is used to design usable species identification primers using Scheme 1 or Scheme 2; the species identification verification unit is used to perform interspecies verification and host verification on the usable species identification primers; and the first evaluation unit is used to evaluate each pair of usable species identification primers.
[0183] The second design evaluation module includes: a second design unit, a subtype identification verification unit, and a second evaluation unit. The second design unit is used to design primers for subtype identification using Scheme 1 or Scheme 2; the subtype identification verification unit is used to perform inter-subtype verification and host verification on the primers for subtype identification; and the second evaluation unit is used to evaluate each pair of primers for subtype identification.
[0184] When the system for designing and evaluating whole-virus primers is specifically implemented, the design and evaluation of primers can be achieved by referring to the method for designing and evaluating whole-virus primers in any of the above embodiments, and the specific implementation steps will not be repeated here.
[0185] The present invention also provides a storage medium storing a computer program, wherein the storage medium includes instructions. When the computer program is executed by a processor, the steps of the method for designing and evaluating a whole-virus primer are implemented.
[0186] The present invention provides a method, system, and storage medium for designing and evaluating whole-virus primers. The present invention develops a method and system for the joint design and evaluation of primers for viral species identification and viral subtype identification. This method and system allows for the design of corresponding primers for different viral subtypes, resulting in universal and highly accurate primer design for viral subtype identification. The present invention obtains usable primers through primer design and specificity verification of the designed primers. Furthermore, the primers are designed based on conserved regions, reducing the probability of amplification failure. Usable primers for species identification and subtype identification are fully verified, with analysis and evaluation based on intra-species, inter-species, (inter-subtype), and host perspectives. This design makes the present invention suitable for designing viral primers for different application scenarios and with different genomic characteristics. On a whole-virus scale, it maximizes the high specificity and accuracy of candidate viral primers. The primers of the present invention are designed based on conserved regions, reducing the probability of amplification failure. This invention comprehensively considers intraspecies, interspecies, (inter-subtype), and host perspectives, resulting in more comprehensive reference information, improved validation dimensions, and enhanced evaluation accuracy, thereby reducing application risks. It also provides a framework for cross-species primer application and offers specificity verification in different hosts. Based on this approach, primers can be selected based on the host, reducing application risks. This approach can generate a full set of viral primers for all viral genomes across species and subtypes, ensuring (to the greatest extent possible) primer accuracy and a broad spectrum of detectable viruses.
[0187] A specific embodiment of the method of the present invention is given below.
[0188] The implementation process is described using Semliki Forest virus (taxid 11033) as an example.
[0189] (1) Genome sequence screening.
[0190] Four genome sequences of this virus were obtained from the GenBank Assembly and RefSeq Assembly of NCBI, with corresponding Accession IDs: GCA_000860285.1, GCA_002889475.1, GCA_031161755.1, and GCF_000860285.1. GCF_000860285.1 and GCF_000860285.1 are the same genome. Therefore, these three genomes, GCF_000860285.1, GCA_002889475.1, and GCA_031161755.1, were used in the next module.
[0191] (2) Species identification primer design
[0192] The virus has a RefSeq genome sequence, and primers for species identification were designed.
[0193] Primers were designed for the GCF_000860285.1 genome as the representative genome. GCA_002889475.1 and GCA_031161755.1 were designated as neighbor genomes. Table 3 shows the grouping of Semliki Forest virus genomes.
[0194] Table 3
[0195]
[0196] GCF_000860285.1 is 11442 bp in length and was cut into two sequences of 10000 bp in length at a step size of 9500 bp to obtain two sequences, NC_003215_001_1_10000.fa and NC_003215_002_9501_11442.fa.
[0197] Primers were designed for these two sequences using Primer3 software. Key parameters included: primer length 18–23 bp, product length 100–300 bp, RCR annealing temperature 59°C, and GC content 30%–70%. The number of primers returned was the integer obtained by rounding off the value of "Template sequence length (bp) / 50." For example, the NC_003215_001_1_10000.fa setting returned 100 primer pairs, and the NC_003215_002_9501_11442.fa setting returned 39 primer pairs.
[0198] Because the two sequences share a 500-bp duplication region, seven identical primer pairs were predicted for this region. After removing duplicates, a total of 232 primer pairs were obtained. Specificity testing was performed on the entire GCF_000860285.1 genome using MFEprimer software (default parameters). The results showed that all primers amplified products and had unique targets.
[0199] Based on the genome sequence and the results of primer3, the information of 232 pairs of primers was extracted. The extracted content and field descriptions are shown in Table 4, which shows the primer information of the representative genome of the virus.
[0200] The specificity of each of the 232 primer pairs was verified on each neighbor's genome using MEFPrimier software (default parameters). Product sequences were aligned using blast software (evalue = 1e-10). Internal code was used to analyze the validation of each primer pair on each neighbor's genome. Using GCA_002889475.1 as an example, some primer validation results are shown in Table 5.
[0201] Table 4
[0202]
[0203] Table 5
[0204]
[0205] Table 6
[0206]
[0207] The pass rate of each primer pair on the neighbors genome was calculated, and the pass rates are shown in Table 6.
[0208] From Table 6 , primers with a pass rate of no less than 50% for the neighbors genome were screened, totaling 230 pairs, which were used as species primers for subsequent analysis.
[0209] A representative viral sequence database was constructed using MEFPrimer software with default parameters. The 230 primer pairs obtained above and the representative viral sequence database were input and the results were output for each genome in the viral library. For each primer pair, the results for product amplification on sequences from non-native species were counted. The results of cross-species validation of the viral library are shown in Table 7 below.
[0210] Table 7
[0211]
[0212] In Table 7, the primer_id column is the name of the primer pair, the target_virus_taxon_num is the number of non-species genomes on which the primer can amplify, and the target_virus_info column is a specific description of the non-species. Species are separated by commas “,”. The information of each species is connected by “|”, which respectively represents the taxid, Organism_name, Species_name and the number of products (hit_num) that the primer can amplify on the species.
[0213] Host Verification: Humans and sheep are the primary vertebrate hosts of this virus. For these two hosts, we used MEFPrimer software with default parameters, inputting the primer set and host genome, to verify that the primers amplified products on the host genomes. The host genome verification results are shown in Table 8 below.
[0214] In Table 8, t_h_n represents target_host_number, which is the number of hosts on which the primer can amplify. The target_host_info column is a detailed description of the host, and the numbers in the brackets represent the number of amplification targets of the primer on the host genome.
[0215] Table 8
[0216]
[0217] The primers were evaluated, and some of the results are shown in Table 9.
[0218] Table 9
[0219]
[0220] Among them, passed_neighbors_number indicates the number of neighbors genomes that have passed, and Level indicates the rating result.
[0221] The statistics of primer numbers at each level are shown in Table 10:
[0222] Table 10
[0223]
[0224] Examples of primers for identifying subtypes of Semliki Forest virus that can be obtained according to the above method are not described here.
[0225] The following are primers for identifying Norovirus subtype H. The specificity of the primers for other subtypes should also be considered (this example analyzes subtypes I, F, J, and G). The process and results of the primers for Norovirus species identification are not shown here.
[0226] For norovirus genome sequence screening, see Table 11.
[0227] Table 11
[0228]
[0229] This embodiment is to design primers for subtypes of norovirus. According to prior knowledge, the viral genome consists of 11 segments, and the difference between different subtypes is mainly the difference in the VP6 gene. In order to enable the primers to avoid the interference of other subtype genomes, the VP6 gene of the target subtype is directly extracted, and then the primers are designed. The target subtype Rotavirus_H has only one genome and does not involve the neighbors genome. The primers designed based on the VP6 gene are specifically verified by the whole genome of Rotavirus_H, and the primers with the only amplification position are retained, totaling 26 pairs of primers. Examples of primers for Rotavirus_H subtype identification are shown in Table 12.
[0230] Non-target subtype genome validation: Using MEFPrimer software with default parameters, the 26 primer pairs obtained above were sequentially validated on four other non-target subtype genomes, for a total of five genomes. The number of amplified targets and subtype information for each primer pair on the non-target subtype genomes were used for subsequent primer evaluation. In this example, none of the 26 primer pairs amplified the non-target subtype genomes.
[0231] Interspecies validation of the virus library: Using MEFPrimer software with default parameters, the 26 primer pairs obtained above and a representative viral sequence database were input. The results were output showing the ability of the primers to amplify products in each genome in the virus library. For each primer pair, the ability to amplify products in sequences not belonging to the same species was counted. The results of the interspecies validation of the Rotavirus_H identification primers are shown in Table 13.
[0232] Host Validation: Since humans are the primary host of norovirus, in this example, primers were validated against the human genome. Using MEFPrimer software with default parameters, the 26 primer pairs obtained above and a representative viral sequence database were input to verify that the primers could amplify products on the host genome. Examples of the results are shown in Table 14 below.
[0233] Table 12
[0234]
[0235] Table 13
[0236]
[0237] Table 14
[0238]
[0239] Primer evaluation. The results of the Rotavirus_H subtype identification primers are shown in Table 15.
[0240] Table 15
[0241]
[0242] Among them, neighbors_nubmer indicates the number of neighbors genomes, and target_other_subtype_num indicates how many non-target subtype genomes the primer can amplify.
[0243] The number of primers at each level was counted, and the results of primer evaluation for Rotavirus_H subtype identification primers are shown in Table 16.
[0244] Table 16
[0245]
[0246] It should be noted that, in the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0247] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or storage media. Thus, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0248] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems and methods according to multiple embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or code, and the part of the module, program segment or code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0249] Although preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they are aware of the basic inventive concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention. Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, the present invention is intended to include such changes and modifications as fall within the scope of the claims and their equivalents.
Claims
1. A method for designing and evaluating whole-virus primers, characterized in that: include: obtaining a viral genome sequence library of the virus; Designing usable primers for species identification and / or designing usable primers for subtype identification; The method of designing the available primers for species identification comprises the following steps: designing the available primers for species identification using scheme 1 or scheme 2, performing interspecies verification and host verification on the available primers for species identification, and evaluating each pair of available primers for species identification; The method of designing primers for subtype identification comprises the following steps: designing primers for subtype identification using Scheme 1 or Scheme 2, performing non-target virus subtype verification, interspecies verification, and host verification on the primers for subtype identification, and evaluating each pair of primers for subtype identification; The first scheme is: determining a representative genome of the virus in a viral genome sequence library, designing primers for the representative genome, using non-representative genomes in the viral genome sequence library to perform specificity verification on the primers for the representative genome, and selecting the primers for the representative genome that meet the verification pass rate as usable primers; The second scheme is: performing multiple sequence alignment on the genome sequences in the viral genome sequence library to identify the conserved regions of the virus, designing primers for the conserved regions, using the viral genome sequence library to perform specificity verification on the primers for the conserved regions, and using the conserved regions primers with a verification pass rate that meets the standard as usable primers; The primers available for species identification are used for virus species verification and host verification, including: Amplifying the available primers for species identification in sequences in a representative viral sequence database, and statistically analyzing the amplification results to obtain viral species verification results of the available primers for species identification, wherein the representative viral sequence database includes representative sequences of multiple viral species; The primers available for species identification are amplified on all host genomes in a host genome database, and the amplification results are statistically analyzed to obtain host verification results of the primers available for species identification, wherein the host genome database includes genomes of several hosts of the virus; The subtype identification primers used for non-target virus subtype verification, virus species verification and host verification specifically include: The available primers for subtype identification are amplified on the genome of the non-target subtype, and the amplification results are statistically analyzed to obtain the non-target virus subtype verification results; The available primers for subtype identification are amplified on all representative sequences in a representative viral sequence database, and the amplification results are statistically analyzed to obtain viral species validation results of the available primers for subtype identification, wherein the representative viral sequence database includes representative sequences of multiple viral species; The primers available for subtype identification are amplified on all host genomes in a host genome database, and the amplification results are statistically analyzed to obtain host verification results of the primers available for subtype identification, wherein the host genome database includes genomes of several hosts of the virus; The statistical amplification situation refers to the amplification effect on the verified genome and the uniqueness of the amplification position.
2. The method for designing and evaluating a whole-virus primer according to claim 1, wherein: The viral genome sequence library is a processed viral genome sequence library, and the viral genome sequence library of the virus is obtained specifically by: obtaining the virus, collecting the subtype information, genome sequence and genome sequence metadata of the virus genome to form an original viral genome sequence library, and performing sequence deduplication and filtering on the original viral genome sequence library to obtain the processed viral genome sequence library.
3. The method for designing and evaluating a whole-virus primer according to claim 1, wherein: If the virus is a single genome virus, use scheme 1 to design primers for species identification and subtype identification; If the virus is a non-monogenome virus, the available primers for species identification designed using Scheme 1 or Scheme 2 are determined based on the variability of the non-monogenome virus and / or the number of genomes of the non-monogenome virus. The available primers for subtype identification designed using Scheme 1 or Scheme 2 are determined based on the variability of the non-monogenome virus and / or the number of genomes of the non-monogenome virus.
4. The method for designing and evaluating a whole-virus primer according to claim 1, wherein: The specificity verification specifically comprises: performing specificity analysis on the primers representing the genome using the verification genome, determining whether the primers representing the genome meet the specificity standard to obtain a specificity verification result, and calculating the specificity verification pass rate of the primers representing the genome according to the specificity verification result; In the first scheme, the verification genome is a non-representative genome in the viral genome sequence library. If the virus does not have a non-representative genome, the primers representing the genome are considered to have passed specific verification, and the verification pass rate is 100%. If the virus has a non-representative genome, the non-representative genome in the viral genome sequence library is used as the verification genome; In the second scheme, the verification genome is the genome of the viral genome sequence library.
5. The method for designing and evaluating a whole-virus primer according to claim 4, wherein: The specificity standard is to meet all of the following conditions: Condition 1: The primers can amplify the genome and the amplification position is unique; Condition 2: The gene amplified by the primers on the verification genome is consistent with the gene amplified by the primers on the representative genome. If the verification genome is a segment genome, the amplification product of the primers on the verification genome and the amplification product of the primers on the verification genome are located in the same segment; Condition 3: Verify that the length of the amplified product on the genome and the length variation of the amplified product of the primer on the representative genome are within the set length variation threshold; Condition 4: Verify that the sequence similarity between the amplified product on the genome and the amplified product of the primer on the representative genome reaches the set sequence similarity threshold condition.
6. The method for designing and evaluating a whole-virus primer according to claim 1, wherein: The evaluation of each pair of available primers for species identification is specifically as follows: for each pair of available primers for species identification, a quality rating is performed based on its verification pass rate, its virus species verification result, and its host verification result; The evaluation of each pair of available primers for subtype identification is specifically as follows: for each pair of available primers for subtype identification, quality rating is performed based on its verification pass rate, its non-target virus subtype verification result, its virus species verification result and its host verification result.
7. A design and evaluation system for whole-virus primers, characterized in that: include: The first design evaluation module is used to design usable primers for species identification using Scheme 1 or Scheme 2, perform interspecies verification and host verification on the usable primers for species identification, and evaluate each pair of usable primers for species identification; The second design evaluation module is used to design primers for subtype identification using Scheme 1 or Scheme 2, perform non-target virus subtype verification, interspecies verification, and host verification on the primers for subtype identification, and evaluate each pair of primers for subtype identification; The first scheme is: determining a representative genome of the virus in a viral genome sequence library, designing primers for the representative genome, using non-representative genomes in the viral genome sequence library to perform specificity verification on the primers for the representative genome, and selecting the primers for the representative genome that meet the verification pass rate as usable primers; The second scheme is: performing multiple sequence alignment on the genome sequences in the viral genome sequence library to identify the conserved regions of the virus, designing primers for the conserved regions, using the viral genome sequence library to perform specificity verification on the primers for the conserved regions, and using the conserved regions primers with a verification pass rate that meets the standard as usable primers; The primers available for species identification are used for virus species verification and host verification, including: Amplifying the available primers for species identification in sequences in a representative viral sequence database, and statistically analyzing the amplification results to obtain viral species verification results of the available primers for species identification, wherein the representative viral sequence database includes representative sequences of multiple viral species; The primers available for species identification are amplified on all host genomes in a host genome database, and the amplification results are statistically analyzed to obtain host verification results of the primers available for species identification, wherein the host genome database includes genomes of several hosts of the virus; The subtype identification primers used for non-target virus subtype verification, virus species verification and host verification specifically include: The available primers for subtype identification are amplified on the genome of the non-target subtype, and the amplification results are statistically analyzed to obtain the non-target virus subtype verification results; The available primers for subtype identification are amplified on all representative sequences in a representative viral sequence database, and the amplification results are statistically analyzed to obtain viral species validation results of the available primers for subtype identification, wherein the representative viral sequence database includes representative sequences of multiple viral species; The primers available for subtype identification are amplified on all host genomes in a host genome database, and the amplification results are statistically analyzed to obtain host verification results of the primers available for subtype identification, wherein the host genome database includes genomes of several hosts of the virus; The statistical amplification situation refers to the amplification effect on the verified genome and the uniqueness of the amplification position.
8. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method for designing and evaluating a whole-virus primer according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Design method and system of targeted pathogenic microorganism sequencing primer
CN118762752A
Oligonucleotide probes for specific identification of noroviruses and other pathogens
US20160034636A1