Method and system for designing and evaluating totivirus primer and storage medium
By designing and verifying viral primers, the problems of insufficient universality, low accuracy and high application risk for virus subtype identification and cross-species applications in the prior art are solved, and high specificity and accuracy of viral primer design is achieved, reducing application risk.
Patent Information
- Application Number
- CN202510436102.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-09
AI Technical Summary
Existing viral primer design methods have problems such as insufficient universality, low accuracy and high application risk when identifying virus subtypes and applying across species.
By obtaining the viral genome sequence library of the virus, use scheme one or two to design the available primers for species identification and subtype identification, and conduct interspecies verification, host verification and specific verification, and screen out primers with validation pass rate as available primers.
Accurate identification of different virus subtypes is achieved, the specificity and accuracy of primers are improved, and the application risks are reduced, so that primers are highly specific and widely applicable at the entire virus scale.
Smart Images

Figure CN119932234A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of primer design, and in particular to a method, system and storage medium for designing and evaluating whole virus primers. Background Art
[0002] Viruses, as a common type of pathogen, have a significant impact on human health and the agricultural output of animals and plants. Timely and accurate identification of virus types and virus subtypes is conducive to the formulation of precise prevention and control strategies. Traditional single virus identification or multi-virus simultaneous identification primers often cannot effectively meet the needs of overall virus spectrum monitoring. In recent years, targeted high-throughput sequencing (tNGS) can support the simultaneous detection of more than 40 common clinical pathogenic viruses ([Almas, S.; Carpenter, RE; Singh, A.; Rowan, C.; Tamrakar, VK; Sharma, R. Deciphering Microbiota of Acute Upper Respiratory Infections: A Comparative Analysis of PCR and mNGSMethods for Lower Respiratory Trafficking Potential. Adv. Respir. Med. 2023,91, 49-65. ]), but there are limitations such as false positives caused by amplification contamination due to poor primer specificity and insufficient virus subtype identification. High specificity and efficiency are important characteristics of high-quality primers and the goals of primer design.
[0003] Usually, when designing primers for pathogenic viruses, the conserved regions of the virus are first calculated using software, then primers are designed for this region, and finally the primer specificity index is calculated. Taking the mainstream application Primer-BLAST software as an example, when designing primers for a single template, its length is limited to 50kb. The software can also design primers for a group of sequences. The specific operation process is: first, the longest sequence is blasted with other sequences to find the conserved regions, then the conserved regions on the longest sequence are extracted and used as representative template sequences, and then primer3 is called for primer design. In addition, primer specificity tests can specify genome databases (such as RefSeq representative genomes) or specific species genomes (such as Home sapiens) to calculate nonspecific amplification efficiency.
[0004] The existing viral primer design methods can basically meet the needs of single virus or multiple virus detection, but they also have certain limitations. The details are as follows:
[0005] (1) Primer design for virus subtype identification is not universal: Each virus subtype has its own particularity and requires personalized primer design. Currently, there is no method or system to design primers for different subtypes of a virus.
[0006] (2) Low accuracy in primer design: There is a lack of comprehensive consideration of the variation regions between different strains of the same species, which means that when primers are designed in these variation regions, amplification failure is very likely to occur for other strains. Since viruses mutate and evolve very rapidly, primers are currently designed only for a specific virus, and their reference information is incomplete, resulting in a high risk of application of the calculated primers. The specificity test of the designed primers is not optimal, which limits the application of primers across species. The unclear specificity in the host makes the risk of nonspecific amplification unknown during application, which in turn makes the application risk high.
[0007] Therefore, there is an urgent need to develop a method and system for the design and evaluation of whole-virus primers. Summary of the invention
[0008] In view of the above problems, the present invention provides a method, system and storage medium for designing and evaluating whole virus primers.
[0009] The technical solution adopted by the present invention to solve the technical problem is as follows:
[0010] In a first aspect, the present invention provides a method for designing and evaluating a whole virus primer, comprising:
[0011] obtaining a viral genome sequence library of the virus;
[0012] Use Scheme 1 or 2 to design usable primers for species identification, perform species verification and host verification on the usable primers for species identification, and evaluate each pair of usable primers for species identification;
[0013] Use Scheme 1 or 2 to design usable primers for subtype identification, perform non-target virus subtype verification, inter-species verification, and host verification on the usable primers for subtype identification, and evaluate each pair of usable primers for subtype identification.
[0014] The first scheme is: determining a representative genome of the virus in a viral genome sequence library, designing primers for the representative genome, using a non-representative genome in the viral genome sequence library to specifically verify the primers for the representative genome, and using the primers for the representative genome that meet the verification pass rate as available primers;
[0015] The second scheme is: perform multiple sequence alignment on the genome sequences in the viral genome sequence library to identify the conserved regions of the virus, design primers for the conserved regions, use the viral genome sequence library to perform specific verification on the primers in the conserved regions, and use the primers in the conserved regions with a verification pass rate that meets the standard as available primers.
[0016] In a preferred embodiment, the viral genome sequence library is a processed viral genome sequence library, and the viral genome sequence library of the virus is obtained specifically by: obtaining the virus, collecting the subtype information, genome sequence and genome sequence metadata of the virus genome to form an original viral genome sequence library, and performing sequence deduplication and filtering processing on the original viral genome sequence library to obtain a processed viral genome sequence library.
[0017] In a preferred embodiment, if the virus is a single-genome virus, Scheme 1 is used to design available primers for species identification and available primers for subtype identification; if the virus is a non-single-genome virus, Scheme 1 or Scheme 2 is used to design available primers for species identification based on the variability of the non-single-genome virus and / or the number of genomes of the non-single-genome virus, and Scheme 1 or Scheme 2 is used to design available primers for subtype identification based on the variability of the non-single-genome virus and / or the number of genomes of the non-single-genome virus.
[0018] In a preferred embodiment, the specificity verification specifically includes: using the verification genome to perform specificity analysis on the primers representing the genome, determining whether the primers representing the genome meet the specificity standard to obtain a specificity verification result, and calculating the specificity verification pass rate of the primers representing the genome according to the specificity verification result;
[0019] In the scheme 1, the verification genome is a non-representative genome in the viral genome sequence library. If the virus does not have a non-representative genome, the primers representing the genome are deemed to have passed specific verification, and the verification pass rate is 100%. If the virus has a non-representative genome, the non-representative genome in the viral genome sequence library is used as the verification genome;
[0020] In the second scheme, the verified genome is the genome of the viral genome sequence library.
[0021] In a preferred embodiment, the specificity criterion is to satisfy all of the following conditions:
[0022] Condition 1: The primers can have an amplification effect on the verified genome and the amplification position is unique;
[0023] Condition 2: The gene amplified by the primers on the verification genome is consistent with the gene amplified by the primers on the representative genome. If the verification genome is a segment genome, the amplification product of the primers on the verification genome and the amplification product of the primers on the verification genome are located in the same segment;
[0024] Condition 3: Verify that the length of the amplified product on the genome and the length variation of the amplified product of the primer on the representative genome are within the set length variation threshold range;
[0025] Condition 4: Verify that the sequence similarity between the amplified product on the genome and the amplified product of the primer on the representative genome reaches the set sequence similarity threshold condition.
[0026] In a preferred embodiment, the inter-species and host validation of the primers available for species identification comprises:
[0027] Amplification of primers available for species identification in sequences in a representative viral sequence database, and statistical amplification to obtain viral species verification results of primers available for species identification, wherein the representative viral sequence database includes representative sequences of multiple viral species,
[0028] Primers available for species identification are amplified on all host genomes in a host genome database, and the amplification conditions are statistically analyzed to obtain host verification results of the primers available for species identification, wherein the host genome database includes genomes of several hosts of the virus.
[0029] In a preferred embodiment, the non-target virus subtype verification, virus species verification and host verification of the available primers for subtype identification specifically include:
[0030] The available primers for subtype identification are amplified on the genome of the non-target subtype, and the amplification situation is statistically analyzed to obtain the verification result of the non-target virus subtype;
[0031] The available primers for subtype identification are amplified on all representative sequences of a representative viral sequence database, and the amplification status is statistically analyzed to obtain the viral species verification results of the available primers for subtype identification, wherein the representative viral sequence database includes representative sequences of multiple viral species;
[0032] The available primers for subtype identification are amplified on all host genomes in the host genome database, and the amplification conditions are statistically analyzed to obtain host verification results of the available primers for subtype identification. The host genome database includes genomes of several hosts of the virus.
[0033] In a preferred embodiment, the evaluation of each pair of available primers for species identification is specifically: for each pair of available primers for species identification, quality rating is performed based on its verification pass rate, its virus species verification result and its host verification result;
[0034] The evaluation of each pair of available primers for subtype identification is specifically as follows: for each pair of available primers for subtype identification, quality rating is performed based on its verification pass rate, its non-target virus subtype verification result, its virus species verification result and its host verification result.
[0035] In a second aspect, the present invention provides a system for designing and evaluating whole-virus primers, comprising:
[0036] The first design evaluation module is used to design usable primers for species identification using Scheme 1 or Scheme 2, perform species verification and host verification on the usable primers for species identification, and evaluate each pair of usable primers for species identification;
[0037] The second design evaluation module is used to design usable primers for subtype identification using Scheme 1 or Scheme 2, perform non-target virus subtype verification, interspecies verification, and host verification on the usable primers for subtype identification, and evaluate each pair of usable primers for subtype identification;
[0038] The first scheme is: determining a representative genome of the virus in a viral genome sequence library, designing primers for the representative genome, using a non-representative genome in the viral genome sequence library to specifically verify the primers for the representative genome, and using the primers for the representative genome that meet the verification pass rate as available primers;
[0039] The second scheme is: perform multiple sequence alignment on the genome sequences in the viral genome sequence library to identify the conserved regions of the virus, design primers for the conserved regions, use the viral genome sequence library to perform specific verification on the primers in the conserved regions, and use the primers in the conserved regions with a verification pass rate that meets the standard as available primers.
[0040] In a third aspect, the present invention provides a storage medium storing a computer program, which, when executed by a processor, implements the steps of the method for designing and evaluating a whole-virus primer described in the first aspect.
[0041] The present invention is a design and evaluation method and system for the joint development of primers for viral species identification and primers for viral subtype identification. The present invention obtains usable primers by designing primers and verifying the specificity of the designed primers, including designing corresponding primers for different subtypes of the virus, and comprehensively verifying the usable primers for species identification and subtype identification, that is, analysis is given based on the perspectives of intra-species, inter-species, (inter-subtype,) host, and the primers are evaluated. Such a design makes the present invention suitable for the design of viral primers in different application scenarios and different genomic characteristics, and has universality and high accuracy for the design of primers for viral subtype identification. On the scale of the whole virus, it ensures the high specificity of the candidate primers for the virus, improves the accuracy of the primers, and reduces the application risk of the primers. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 A flowchart of a method for designing and evaluating whole-virus primers;
[0043] Figure 2 Flow chart of the design and evaluation of usable primers for viral species identification;
[0044] Figure 3 Flow chart for designing primers that can be used for viral species identification using Scheme 1;
[0045] Figure 4 A validation flow chart for obtaining usable primers for species identification primers using the protocol;
[0046] Figure 5 Flow chart for designing primers that can be used for viral species identification using scheme 2;
[0047] Figure 6 Flowchart for the design and evaluation of useful primers for viral subtype identification;
[0048] Figure 7 A systematic framework diagram for the design and evaluation of whole-virus primers. DETAILED DESCRIPTION
[0049] In order to more clearly understand the above-mentioned objects, features and advantages of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0050] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited to the specific embodiments disclosed below.
[0051] It should be noted that the descriptions involving "first", "second", etc. in this application are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the ability of ordinary technicians in this field to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such combination of technical solutions does not exist and is not within the scope of protection required by this application.
[0052] The existing methods and systems for designing viral primers not only lack universal applicability, but also have low accuracy and high application risks. To this end, this application provides a method for designing and evaluating all-viral primers, such as Figure 1 , the method comprising:
[0053] obtaining a viral genome sequence library of the virus;
[0054] Design usable primers for species identification and / or design usable primers for subtype identification.
[0055] The method of designing the available primers for species identification is as follows: using scheme 1 or scheme 2 to design the available primers for species identification, performing species verification and host verification on the available primers for species identification, and evaluating each pair of available primers for species identification;
[0056] The method of designing available primers for subtype identification is as follows: using scheme one or scheme two to design available primers for subtype identification, performing non-target virus subtype verification, species verification and host verification on the available primers for subtype identification, and evaluating each pair of available primers for subtype identification.
[0057] The first scheme is: determining the representative genome of the virus in the viral genome sequence library, designing primers for the representative genome, using the non-representative genome in the viral genome sequence library to perform specificity verification on the primers of the representative genome, and using the primers of the representative genome that meet the specificity verification pass rate as available primers.
[0058] The second scheme is: perform multiple sequence alignment on the genome sequences in the viral genome sequence library to identify the conserved regions of the virus, design primers for the conserved regions, use the viral genome sequence library to perform specificity verification on the primers in the conserved regions, and use the primers in the conserved regions with a specificity verification pass rate that meets the standard as available primers.
[0059] Here, there is no restriction on the order of the primers available for species identification and the primers available for subtype identification.
[0060] In this embodiment, the viral genome sequence library is a processed viral genome sequence library, and the viral genome sequence library of the virus is obtained specifically by: obtaining the virus, collecting the subtype information, genome sequence and genome sequence metadata of the virus genome to form an original viral genome sequence library, and performing sequence deduplication and filtering processing on the original viral genome sequence library to obtain the processed viral genome sequence library.
[0061] The method of this embodiment is divided into two parts: design and evaluation of primers for viral species identification and design and evaluation of primers for viral subtype identification. The order of the two parts is not limited.
[0062] Part I. Design and evaluation of primers for viral species identification. Figure 2 , including the following steps S11~S15.
[0063] S11. Obtain a virus (a virus for which primers are to be designed), collect the genome sequence and genome sequence metadata of the virus for which primers are to be designed to form an original virus genome sequence library 1, and screen the genome sequence in the original virus genome sequence library 1 to obtain a processed virus genome sequence library 1.
[0064] The screening of the genome sequence is to perform sequence deduplication and filtering on the original viral genome sequence library.
[0065] Specifically, all genome sequences and metadata of the virus are collected from public databases to form an original viral genome sequence library 1, which is then deduplicated and low-quality sequence filtered, so that the original viral genome sequence library 1 can retain valid information while streamlining the amount of sequence data, thereby improving primer design efficiency.
[0066] As an embodiment, the sequence deduplication is specifically to retain only one 100% identical genome sequence through cd-hit software. The sequence filtering is to remove short fragment sequences, and the specific length threshold of the filtered out can be designed according to the actual situation.
[0067] S12. Design and verify primers for virus species identification to obtain usable primers for species identification.
[0068] S12.1. Determine whether the virus species identification primer design is S12.2 or S12.3 based on whether it is a single genome virus. The specific determination is also based on the variability of non-single genome viruses and / or the number of genomes of non-single genome viruses. The variability described herein can be understood as a mutation rate.
[0069] The present invention provides two schemes for designing primers for virus species identification. For single-genome viruses, scheme one (S12.2) is directly selected to design species identification primers. Since it does not involve multiple sequence alignment, the time cost and calculation cost are lower than those of scheme two. For non-single-genome viruses, scheme two (S12.3) can be directly selected to design primers. The conserved region is first extracted and then the species identification primers are designed. In this way, primers with high species specificity can be obtained more quickly. However, it is not limited to non-single-genome viruses that all must select scheme two.
[0070] Typically, a variability threshold and a genome number threshold may be set to determine whether to select Scheme 1 or Scheme 2 for a non-single-genome virus based on the variability threshold and / or genome number threshold of the non-single-genome virus.
[0071] In one embodiment, when the variability of a non-single genome virus exceeds a variability threshold and the number of genomes exceeds a genome number threshold, option 2 is selected, otherwise option 1 is selected; in one embodiment, when at least one of the variability of a non-single genome virus exceeds a variability threshold and the number of genomes exceeds a genome number threshold is true, option 2 is selected, otherwise option 1 is selected; in one embodiment, variability or the number of genomes is not considered, for example, variability is not considered, when the number of genomes of a non-single genome virus exceeds a genome number threshold, option 2 is selected, otherwise option 1 is selected, for example, the number of genomes is not considered, when the variability of a non-single genome virus exceeds a variability threshold, option 2 is selected, otherwise option 1 is selected. It can be understood that the above embodiments are only examples and not limitations.
[0072] As a specific example, for viruses whose genome length is less than 100 kb, the genome number threshold is a number less than or equal to 50. If the genome length of the virus for which primers are to be designed exceeds 100 kb, the genome number threshold is a number less than or equal to 10.
[0073] S12.2. Determine the representative genome in the processed viral genome sequence library one, design species identification primers (primers for species identification) for the representative genome, and obtain a species identification primer set for the representative genome; use other genomes in the processed viral genome sequence library one (other genomes are non-representative genomes) to verify the species identification primers in the species identification primer set of the representative genome, screen the species identification primers of the representative genome with a passing rate that meets the verification standard as the available primers for species identification, obtain the available primer set for species identification, and proceed to S13.
[0074] The process of S12.2 can be found in Figure 3 .
[0075] Determine the representative genome: Select a genome from the processed viral genome sequence library 1 as the representative genome. The selection principle can be targeted strain screening based on the scientific problem being studied, or it can be preferential selection of RefSeq (Reference Sequence) genomes, etc. In the processed viral genome sequence library 1, in addition to the representative genome, the other genomes in the processed viral genome sequence library 1 are called neighbors genomes, or non-representative genomes, which are used to verify primers and serve as verification genomes for primer verification.
[0076] Use primer3 software to design species identification primers for representative genomes: If the length of all sequences of the representative genome is within 10kb, primers can be designed directly through primer3 software. However, when the length of the representative genome exceeds 10kb or the length of the contigs (or segments) of the representative genome exceeds 10kb, due to the length limit of primer3 software, long sequences (long sequences for primer3 software) need to be cut, and a window of 10kb and a step of 9.5kb are used for cutting, or a mode with a 500bp overlapping window is used for cutting (the product length is set at 100bp-300bp, and the overlapping window is set to 500bp so that primers will not be lost due to sequence cutting). Primers are designed for the cut results (sub-fragments) using primer3 software, and then the primer results are merged. Deduplication is required when merging the primer results.
[0077] Keep the primers that are unique in the amplification position to obtain the species identification primer set representing the genome. When the viral genome sequence (corresponding to the representative genome at this time) is a complete sequence, Primer3 will consider the uniqueness of the amplification position when designing primers, but when the viral genome is composed of multiple segments, Primer3 software will not consider the specificity of primers on different segments, so it is also necessary to screen out primers that are unique in the amplification position on the entire genome. Only such primers can be included in the species identification primer set representing the genome.
[0078] As an example, the main parameters of primer 3 design are: primer length 18bp-23bp, product length 100bp-300bp, RCR annealing temperature 59°C, and GC content range 30%-70%.
[0079] The validation results were obtained by using the neighbors genome to validate the species identification primers: to determine whether there is a neighbors genome. If the virus does not have a neighbors genome, the species identification primer set representing the genome can be directly used as the available primer set for species identification, that is, the primers of the species identification primer set representing the genome are directly regarded as having passed the specificity validation, and the validation pass rate is 100%; if the virus has a neighbors genome, the species identification primer set representing the genome is specifically tested on each neighbors genome ( Figure 3 In the , "in-house script" means "internal script" ).
[0080] The flow chart of this step is as follows Figure 4 As shown in , each neighbor genome is a validation genome. Figure 4 As shown, MEFprimer software is used to analyze the specificity of the primers to be verified on the target genome (i.e., the neighbors genome at this time), and then the specificity is determined to determine whether it meets the specificity standard to obtain the specificity verification result. The specificity verification pass rate is calculated based on the specificity verification result. Determine whether the specificity meets the following conditions: (1) The species identification primer can have an amplification effect on the verification genome and the amplification position is unique (if there is no amplification, determine whether the genome is a segment genome and whether it has the same segment as the primer template. The purpose of this operation is to finely mark the reasons for failure of the verification); (2) The gene amplified by the species identification primer on the verification genome is consistent with the gene amplified by the species identification primer on the representative genome. If the verification genome is a segment genome, it is also required that the amplification product of the primer on the verification genome and the amplification product of the species identification primer on the verification genome are located in the same segment; (3) the length of the amplified product on the verification genome and the length of the amplified product of the species identification primer on the representative genome vary within the set length variation threshold (e.g., 10%); (4) the sequence similarity of the amplified product on the verification genome and the amplified product of the species identification primer on the representative genome reaches the set sequence similarity threshold condition (e.g., blast comparison results show identity ≥ 95%, coverage ≥ 95%, identity is the ratio of the consistent base sites in the two sequences to the total number of bases; coverage is the ratio of the length of the comparison region to the total sequence length). If all of the above conditions are met, it is considered to meet the specificity standard, and the specific verification result of the primer pair on a verification genome is verified. Perform specificity analysis and judgment one by one to obtain a verified representative genome species identification primer set.
[0081] The species identification primer set of the representative genome that has passed the verification is screened for virus species identification primers to obtain the available primer set for species identification. The available primers are screened by setting the pass rate threshold. Specifically: After the neighbors genome is verified, according to the principle of "rounding up and leaving even numbers" (optional), the pass rate of each pair of species identification primers representing the genome relative to all verified genomes is calculated to pass the specificity standard, which is called the verification pass rate. Then, the species identification primers representing the genome that meet the pass rate threshold (such as 0.5) are screened as available primers for species identification, forming the available primer set for species identification.
[0082] S12.3, see Figure 5 , perform multiple sequence alignment on the genome sequences in the processed viral genome sequence library one, the number of genome sequences in the processed viral genome sequence library one needs to be greater than or equal to 2, identify the conserved region of the virus according to the multiple sequence alignment result, design species identification primers for the conserved region, obtain a species identification primer set for the conserved region, use the genome in the processed viral genome sequence library one to verify the primers in the species identification primer set for the conserved region, screen the species identification primers in the conserved region with a verification pass rate that meets the standard as available primers for species identification, obtain a set of available primers for species identification, and perform S13.
[0083] Multiple sequence alignment: It can be part or all of the genome sequences in the processed viral genome sequence library 1, and multiple sequence alignment is performed using Clustal OMEGA software (with default parameters) to obtain. If the genome is a segment genome, multiple sequence alignment is performed based on the "same segment fragment" or "same gene" of different genomes.
[0084] Conserved region identification: Gblocks software was used, and its important parameters were set as follows:
[0085] -b1 The minimum support number of single bases in the conservative region is set to the minimum integer that is not less than 50% of the genome number;
[0086] -b2 The minimum support number of single bases at the boundary of the conservative region is set to the minimum integer that is not less than 85% of the number of genomes;
[0087] -b3 The maximum number of consecutive non-conservative bases is set to 8bp;
[0088] -b4 The minimum length of the conserved region is set to 200 bp;
[0089] -b5 Whether to allow gap positions (None, With Half, All). Set to a to allow.
[0090] Extract the conserved region sequence as the template sequence: traverse each base position in the conserved region, and retain the base with the most support numbers at each position (if two bases have the same support number, the bases are prioritized in alphabetical order). If the base with the most support number is "-", skip it.
[0091] According to the extracted conserved region sequence, primers were designed for the conserved region using primer3 software. The specific process can refer to the design of primers for representative genomic species identification, except that the input file was changed from the representative genomic sequence to the conserved region sequence.
[0092] The species identification primers were verified using the genome of the processed viral genome sequence library one to obtain a verification result: the species identification primer set in the conserved region was specifically tested on each genome of the processed viral genome sequence library one (ie, specific verification).
[0093] As an embodiment, each genome in the processed viral genome sequence library 1 is a verification genome. As another embodiment, except for the genome with the template sequence in the processed viral genome sequence library 1, all are used as verification genomes and can be considered as neighbors genomes. MEFprimer software was used to analyze the specificity of the species identification primers on the target genome (verification genome), and then the results were tested to see if they met the following conditions: (1) the species identification primers had an amplification effect on the verification genome and the amplification position was unique; (2) the gene amplified by the species identification primers on the verification genome was consistent with the gene amplified by the species identification primers in the conserved region. If the verification genome was a segment genome, the amplification product of the primers on the verification genome and the amplification product of the species identification primers on the verification genome were also required to be located in the same segment; (3) the length of the amplification product on the verification genome and the length of the amplification product of the species identification primers in the conserved region varied within the set threshold range (e.g., 10%); (4) the sequence similarity between the amplification product on the verification genome and the amplification product of the species identification primers in the conserved region reached the set threshold condition (e.g., blast comparison results showed identity ≥ 95%, coverage ≥ 95%, identity was the ratio of the consistent base sites in the two sequences to the total number of bases; coverage was the ratio of the length of the comparison region to the total sequence length). If all the above conditions are met, the test is considered to have passed, and a set of species identification primers for the verified conserved region is obtained.
[0094] The species identification primer sets of the verified conserved regions are screened to obtain the available primer sets for species identification: After the genome verification is completed, according to the principle of "rounding off and leaving even numbers", the passing rate of each pair of species identification primers in the conserved regions that passes the specificity standard relative to all verified genomes is calculated (statistical), which is called the verification passing rate. Then, the species identification primers in the conserved regions that meet the passing rate threshold (for example, 0.5) are screened as the available primers for species identification, forming the available primer set for species identification.
[0095] S13. The available primer set for species identification is amplified on all representative sequences in the representative virus sequence database to obtain the amplification status of the available primers for species identification in the sequences in the representative virus sequence database, and the amplification status (the number of amplified targets and species information of each pair of available primers for species identification in the genomes of other species) is counted to obtain the verification results between virus species. The representative virus sequence database includes representative sequences of multiple virus species.
[0096] S13.1. Construct a representative virus sequence database.
[0097] For each virus species, all its sequences are collected and then de-redundanted, and representative sequences are retained. The representative sequences of each species constitute a representative virus sequence database.
[0098] The process of building a representative viral sequence database includes three key steps:
[0099] ① Metadata cleaning and standardization: Standardize metadata such as host, country, time, and separation source. Delete duplicate sequences of metadata.
[0100] ② Sequence clustering and quality control: MMseq2 was used to cluster sequences based on species, and sequences that were 100% covered by other fragments were deleted. BlastN was used to perform pairwise sequence alignment and remove sequences that could not be clustered normally.
[0101] ③ Representative sequence screening: retain the sequences with the most complete metadata and sequence information for each branch. Such sequences are used as representative sequences. The representative sequences of each viral species constitute the representative viral sequence database.
[0102] S13.2. Perform interspecies verification on the primers available for species identification to obtain interspecies verification results of the virus.
[0103] Use MEFPrimer software with default parameters, input the available primer set for species identification and the representative viral sequence database, and the MEFPrimer software will output the amplification status of the available primers for species identification in the sequences of the representative viral sequence database, count the number of amplified targets and species information of each pair of available primers for species identification on the genomes of other species, and obtain the viral species verification result, called viral species verification result 1, or the viral species verification result of the available primers for species identification, which is used for subsequent primer evaluation.
[0104] S14. Amplify the available primer set for species identification on all host genomes in the host genome database, count the amplification status, and obtain the host verification result. The host genome database includes the genomes of several hosts of the virus (corresponding to the available primers for species identification).
[0105] Specifically, S14 includes:
[0106] S14.1. Construct a host genome database of the virus.
[0107] For each virus, search for its corresponding common host through DNA sequence databases (such as NCBI) or literature. For common hosts, you can define what is common. For example, among all known hosts of a virus, if the proportion of a certain host exceeds 10% of the total number of hosts, it is defined as common. Download the genome sequences of common hosts to form a host genome database. Each virus has a host genome database.
[0108] S14.2. Perform host validation on primers available for species identification.
[0109] Use MEFPrimer software with default parameters, input the available primer set for species identification corresponding to the virus and the host genome database corresponding to the virus, and the software will output the amplification of the available primers for species identification in the host genome. Count the number of amplified targets in the host of each pair of available primers for species identification and the host species information to obtain the host verification results for subsequent primer evaluation.
[0110] It can be understood that the order of S13 and S14 is not limited.
[0111] S15. Perform a quality rating on each pair of available primers for species identification in the set of available primers for species identification. Specifically, perform a quality rating on each pair of available primers for species identification based on its (the pair of available primers for species identification) verification genome pass rate (the verification pass rate), its viral species verification result (viral species verification result one), and its host verification result (host verification result one).
[0112] As an example but not a limitation, factor A represents the pass rate of the verified genome (if there is no verified genome A, it is set to 100%), factor B represents the number of other viral species that can be amplified, and factor C represents the number of host genomes that can be amplified. The primer evaluation reference scheme can be referred to Table 1.
[0113] Table 1
[0114] Among them, L1 level is: no verified genome or the verified genome pass rate is 100%, and there is no amplification effect in the genome sequences of other viral species and no amplification effect on the host genome. This type of primer belongs to the most specific species identification primer set.
[0115] Level L2: Although the verification genome pass rate is less than 100%, it is more than 50%, and there is no amplification effect in the genome sequences of other viral species and no amplification effect on the host genome. This type of primer belongs to the species identification primer set with high specificity.
[0116] Level L3-L4: The verification genome pass rate is not less than 50%, and there is an amplification effect on no more than 2 viral genomes of non-species. It is divided into L3 and L4 according to whether there is an amplification effect on the host genome. When using primers of this level, it is necessary to consider the coexistence of contaminating viruses and hosts in the sample to be tested in combination with the application scenario.
[0117] Level L5-L6: The verification genome pass rate is not less than 50%, and there is an amplification effect on more than 2 viral genomes of non-species viruses. It is divided into L5 and L6 according to whether there is an amplification effect on the host genome. When using level primers, they should be used with caution in combination with application scenarios and needs.
[0118] The second part is the design and evaluation scheme of available primers for virus subtype identification. The process framework is shown in the figure below. Figure 6 The design and evaluation of primers for viral subtype identification includes the following steps S21 to S25.
[0119] S21. Obtain a virus (a virus for which primers are to be designed), collect the genome sequence and genome sequence metadata of the virus for which primers are to be designed, and the subtype information of the virus genome to form an original virus genome sequence library 2, and screen the genome sequence in the original virus genome sequence library 2 to obtain a processed virus genome sequence library 2.
[0120] The screening of the genome sequence is to perform sequence deduplication and filtering on the original viral genome sequence library. The specific process can be found in S11.
[0121] It can be understood that the design of available primers for species identification and the design of available primers for subtype identification share the same viral genome sequence library. Here, S11 and S21 can be the same step. In some embodiments, only S21 is performed without S11, that is, the processed viral genome sequence library 2 is used in S12 to S14.
[0122] S22. Extract the genome set of the target subtype, design and verify primers for subtype identification to obtain usable primers for subtype identification.
[0123] If you know which genes or fragments determine different subtypes through literature research or prior knowledge, you can first extract the variant genes or fragments of each genome and then design primers for subtype identification.
[0124] S22.1. Determine whether S22.2 or S22.3 is used for the design of primers for virus subtype identification based on whether the virus is a single genome virus. Specifically, the determination is also based on the variability (mutation rate) of non-single genome viruses and / or the number of genomes of non-single genome viruses.
[0125] For any virus, the design of the available primers for species identification and the design of the primers for subtype identification can adopt the same scheme, or each can have its own judgment criteria when designing. In this embodiment, in order to simplify the steps, if scheme 1 is adopted in S12, then scheme 1 is also adopted in S22, and if scheme 2 is adopted in S12, then scheme 2 is also adopted in S22.
[0126] The present invention provides two schemes for designing primers for virus subtype identification. For single-genome viruses, scheme one (S22.2) is directly selected to design subtype identification primers. Since it does not involve multiple sequence alignment, the time cost and calculation cost are lower than those of scheme two (S22.3). For non-single-genome viruses, S22.3 can be directly selected to design primers. The conserved region is first extracted and then the subtype identification primers are designed. In this way, primers with high specificity can be obtained more quickly. However, it is not limited to that non-single-genome viruses must select S22.3.
[0127] Typically, a variability threshold and a genome number threshold may be set to determine whether a non-single genome virus selects S22.2 or S22.3 based on the variability threshold and / or the genome number threshold.
[0128] In one embodiment, when the variability of a non-single genome virus exceeds a variability threshold and the number of genomes exceeds a genome number threshold, S22.3 is selected, otherwise S22.2 is selected; in one embodiment, when at least one of the variability of a non-single genome virus exceeds a variability threshold and the number of genomes exceeds a genome number threshold is true, S22.3 is selected, otherwise S22.2 is selected; in one embodiment, variability or the number of genomes is not considered, for example, variability is not considered, when the number of genomes of a non-single genome virus exceeds a genome number threshold, S22.3 is selected, otherwise S22.2 is selected, for example, the number of genomes is not considered, when the variability of a non-single genome virus exceeds a variability threshold, S22.3 is selected, otherwise S22.2 is selected. It can be understood that the above embodiments are only examples and not limitations.
[0129] S22.2. Determine the representative genome in the processed viral genome sequence library 2, design primers for subtype identification of the representative genome, and obtain a subtype identification primer set for the representative genome; use other genomes (non-representative genomes) in the processed viral genome sequence library 2 to verify the subtype identification primers in the subtype identification primer set of the representative genome, screen the subtype identification primers of the representative genome with a pass rate that meets the verification standard as usable primers for subtype identification, obtain a usable primer set for subtype identification, and proceed to S23.
[0130] Determine the representative genome: Select a genome from the processed viral genome sequence library 2 as the representative genome. The selection principle can be targeted strain screening based on the scientific problem being studied, or it can be preferential selection of RefSeq (Reference Sequence) genomes, etc. In the processed viral genome sequence library 2, in addition to the representative genome, the other genomes in the processed viral genome sequence library 2 are called neighbors genomes, or non-representative genomes, which are used to verify primers and serve as verification genomes for primer verification.
[0131] Use primer3 software to design subtype identification primers for the representative genome: If the length of all sequences of the representative genome is within 10kb, primers can be designed directly through primer3 software. However, when the length of the representative genome exceeds 10kb or the length of the contigs (or segments) of the representative genome exceeds 10kb, due to the length limit of primer3 software, long sequences (long sequences for primer3 software) need to be cut, and a 500bp overlapping window mode is used for cutting (the product length is set at 100bp-300bp, and the overlapping window is set to 500bp so that primers will not be lost due to sequence cutting). Use primer3 software to design primers for the results after cutting, and then merge the primer results. Deduplication needs to be performed when merging the primer results.
[0132] When the viral genome sequence (corresponding to the representative genome at this time) is a complete sequence, Primer3 will consider the uniqueness of the amplification position when designing primers. However, when the viral genome consists of multiple segments, Primer3 software will not consider the specificity of primers in different segments. Therefore, it is also necessary to screen out primers with unique amplification positions on the entire genome. Only such primers can be included in the subtype identification primer set of the representative genome.
[0133] As an example, the main parameters of primer 3 design are: primer length 18bp-23bp, product length 100bp-300bp, RCR annealing temperature 59°C, and GC content range 30%-70%.
[0134] The subtype identification primers were verified using the neighbors genome and the verification results were obtained: when the virus does not have a neighbors genome, the subtype identification primer set representing the genome can be directly used as an available primer set for subtype identification, that is, the primers of the subtype identification primer set representing the genome are directly regarded as having passed the specificity verification, and the verification pass rate is 100%; if the virus has a neighbors genome, the subtype identification primer set representing the genome will be specifically tested on each neighbors genome.
[0135] Each neighbor genome is a validation genome. MEFprimer software is used to analyze the specificity of primers on the target genome (target validation genome, i.e., the neighbors genome at this time), and then test whether the results meet the following conditions: (1) the subtype identification primers can have an amplification effect on the validation genome and the amplification position is unique; (2) the amplified genes of the subtype identification primers on the validation genome are consistent with the amplified genes of the subtype identification primers on the representative genome. If the validation genome is a segment genome, it is also required that the amplified products of the primers on the validation genome and the amplified products of the subtype identification primers on the validation genome are consistent. The products are located on the same segment; (3) The length of the amplified product on the verification genome and the length of the amplified product of the subtype identification primer on the representative genome vary within the set length variation threshold (e.g., 10%); (4) The sequence similarity of the amplified product on the verification genome and the amplified product of the subtype identification primer on the representative genome reaches the set sequence similarity threshold (e.g., blast comparison results show identity ≥ 95%, coverage ≥ 95%, identity is the ratio of the consistent base sites in the two sequences to the total number of bases; coverage is the ratio of the length of the comparison region to the total sequence length). If all of the above conditions are met, it is considered to meet the specificity standard, and the specificity verification result of the primer pair on a verification genome is verified. Specificity analysis and judgment are performed one by one to obtain a subtype identification primer set for the representative genome that has passed the verification.
[0136] The subtype identification primer set representing the verified genome is subjected to viral subtype identification primer screening to obtain an available primer set for subtype identification: After the neighbors genome verification is completed, according to the principle of "rounding off and leaving even numbers", the pass rate of each pair of subtype identification primers representing the genome relative to all neighbors genomes that passes the specificity standard is calculated, which is called the verification pass rate. Then, subtype identification primers representing the genome that meet the pass rate threshold (for example, 0.5) are screened as available primers for subtype identification, forming an available primer set for subtype identification.
[0137] S22.3. Perform multiple sequence alignment on the genome sequences in the processed viral genome sequence library 2, identify the conserved regions of the virus based on the multiple sequence alignment results, design subtype identification primers for the conserved regions, and obtain a subtype identification primer set for the conserved regions. Use the genome in the processed viral genome sequence library 2 to verify the primers in the subtype identification primer set for the conserved regions, screen the subtype identification primers for the conserved regions with a verification pass rate that meets the standard as available primers for subtype identification, obtain a available primer set for subtype identification, and proceed to S23.
[0138] Multiple sequence alignment: It can be part or all of the genome sequences in the processed viral genome sequence library 2, and multiple sequence alignment is performed using Clustal OMEGA software (with default parameters) to obtain. If the genome is a segment genome, multiple sequence alignment is performed based on the "same segment fragment" or "same gene" of different genomes.
[0139] Conserved region identification: Gblocks software was used, and its important parameters were set as follows:
[0140] -b1 The minimum support number of single bases in the conservative region is set to the minimum integer that is not less than 50% of the genome number;
[0141] -b2 The minimum support number of single bases at the boundary of the conservative region is set to the minimum integer that is not less than 85% of the number of genomes;
[0142] -b3 The maximum number of consecutive non-conservative bases is set to 8bp;
[0143] -b4 The minimum length of the conserved region is set to 200 bp;
[0144] -b5 Whether to allow gap positions (None, With Half, All). Set to a to allow.
[0145] Extract the conserved region sequence as the template sequence: traverse each base position in the conserved region, and retain the base with the most support numbers at each position (if two bases have the same support number, the bases are prioritized in alphabetical order). If the base with the most support number is "-", skip it.
[0146] Primer3 software was used to design primers for the conserved region sequences. The specific process can refer to the design of primers for representative genome subtype identification, except that the input file was changed from the representative genome sequence to the conserved region sequence.
[0147] The subtype identification primers were verified using the processed genomes of the viral genome sequence library II to obtain a verification result: the subtype identification primer set in the conserved region was specifically tested on each genome of the processed viral genome sequence library II.
[0148] As an embodiment, each genome in the processed viral genome sequence library 2 is a verification genome. As another embodiment, except for the genomes with template sequences in the processed viral genome sequence library 2, all are used as verification genomes and can be considered as neighbors genomes. MEFprimer software was used to analyze the specificity of the subtype identification primers on the target genome (validation genome), and then the results were tested to see if they met the following conditions: (1) the subtype identification primers had an amplification effect on the validation genome and the amplification position was unique; (2) the gene amplified by the subtype identification primers on the validation genome and the gene amplified by the subtype identification primers in the conserved region were consistent. If the validation genome was a segment genome, the amplification product of the primers on the validation genome and the amplification product of the subtype identification primers on the validation genome were also required to be located in the same segment; (3) the length of the amplification product on the validation genome and the length of the amplification product of the subtype identification primers in the conserved region varied within the set length variation threshold (e.g., 10%); (4) the sequence similarity between the amplification product on the validation genome and the amplification product of the subtype identification primers in the conserved region met the set sequence similarity threshold (e.g., blast comparison results showed identity ≥ 95%, coverage ≥ 95%, identity was the ratio of the consistent base sites in the two sequences to the total number of bases; coverage was the ratio of the length of the alignment region to the total sequence length). If all the above conditions are met, the test is considered to have passed, and a subtype identification primer set of the verified conserved region is obtained.
[0149] It is understandable that in different specificity verifications, the setting of the relevant threshold may be different.
[0150] The subtype identification primer sets of the verified conserved regions are screened to obtain the available primer sets for subtype identification: After the genome verification is completed, according to the principle of "rounding off and leaving even numbers", the passing rate of each pair of subtype identification primers in the conserved regions that passes the specificity standard relative to all verified genomes is calculated, which is called the verification passing rate. Then, the subtype identification primers in the conserved regions that meet the passing rate threshold (for example, 0.5) are screened as the available primers for subtype identification, forming the available primer set for subtype identification.
[0151] S23. Non-target virus subtype verification: The available primer set for subtype identification is amplified on the genome of the non-target subtype, and the amplification status is counted to obtain the non-target virus subtype verification result.
[0152] Use MEFPrimer software to verify the amplification of the available primer sets for subtype identification on the genome of non-target subtypes. Count the number of amplified target points and subtype information of each pair of subtype identification primers on the genome of non-target subtypes for subsequent primer evaluation. The genome of non-target subtypes comes from the processed viral genome sequence library 2.
[0153] It can be understood that, for example, a virus has a first subtype and a second subtype, then S23 is verification: verifying the amplification of the available primer set for subtype identification of the first subtype on the genome of the second subtype, and verifying the amplification of the available primer set for subtype identification of the second subtype on the genome of the first subtype.
[0154] S24. The available primer set for subtype identification is amplified on all representative sequences in the virus representative sequence database, and the amplification situation (the number of amplified targets and species information of each pair of available primers for subtype identification on the genome of other species) is counted to obtain the verification results between virus species. The virus representative sequence database includes representative sequences of multiple virus species.
[0155] S24.1. Construct a representative virus sequence database.
[0156] For each virus species, all its sequences are collected and then de-redundanted, and its representative sequences are retained. The representative sequences of each virus species constitute a virus representative sequence database. The virus representative sequence database here is exactly the same as the virus representative sequence database in S13.1. The virus representative sequence database obtained in S13.1 can be directly used without re-construction.
[0157] The process of building a representative viral sequence database includes three key steps:
[0158] ① Metadata cleaning and standardization: Standardize metadata such as host, country, time, and separation source. Delete duplicate sequences of metadata.
[0159] ② Sequence clustering and quality control: MMseq2 was used to cluster sequences based on species, and sequences that were 100% covered by other fragments were deleted. BlastN was used to perform pairwise sequence alignment and remove sequences that could not be clustered normally.
[0160] ③ Representative sequence screening: retain the sequences with the most complete metadata and sequence information for each branch. Such sequences are used as representative sequences. The representative sequences of each viral species constitute the representative viral sequence database.
[0161] S24.2. Perform inter-species verification on the available primers for subtype identification to obtain inter-species verification results.
[0162] Use MEFPrimer software with default parameters, input the available primer set for subtype identification and the representative viral sequence database (processed viral genome sequence library 2), and the MEFPrimer software will output the amplification of the subtype identification primers in the sequences of the representative viral sequence database, count the number of amplified targets and species information of each pair of subtype identification primers in the genomes of other species, and obtain the viral species verification results, which are called viral species verification results 2, or viral species verification results of available primers for subtype identification, for subsequent primer evaluation.
[0163] S25. Amplify the available primer set for subtype identification on all host genomes in the host genome database, count the amplification status, and obtain the host verification result. The host genome database includes the genomes of several hosts of the virus (corresponding to the available primers for subtype identification).
[0164] Specifically, S25 includes:
[0165] S25.1. Construct a host genome database of the virus.
[0166] For each virus, search for its corresponding common host through a DNA sequence database (such as NCBI) or literature. For common hosts, you can set what is common by yourself. For example, among all known hosts of a virus, the proportion of a certain host exceeds 10% of the total number of hosts, which is defined as common. Download the genome sequences of common hosts to form a host genome database. Each virus has a host genome database. In this embodiment, for the same virus, the host genome database of S25 is the same as the host genome database of S14, that is, there is no need to repeatedly construct the host genome database.
[0167] S25.2. Perform host validation on the primers available for subtype identification.
[0168] Use MEFPrimer software with default parameters, input the available primer set for subtype identification corresponding to the virus and the host genome database corresponding to the virus, and the software will output the amplification of the available primers for subtype identification in the host genome. Count the number of amplified targets and species information in the host for each pair of available primers for subtype identification, and obtain the host verification results for subsequent primer evaluation.
[0169] It can be understood that the order of S23, S24 and S25 is not limited.
[0170] S26. Perform a quality rating on each pair of available primers for subtype identification in the set of available primers for subtype identification. Specifically, for each pair of available primers for subtype identification, a quality rating is performed based on its (the pair of available primers for subspecies identification) verification genome pass rate, its non-target virus subtype verification result, its virus species verification result (virus species verification result 2) and its host verification result (host verification result 2).
[0171] As an example but not a limitation, factor A represents the pass rate of the verified genome (no verified genome A is set to 100%), factor B represents the number of other viral species that can be amplified, factor C represents the number of host genomes that can be amplified, and factor D represents the number of non-target subtypes that can be amplified. The reference scheme for subtype identification primer evaluation can be referred to Table 2, L1 represents the first level, and L2, L3, etc. can be understood by the same logic.
[0172] Table 2
[0173] It should be understood that the above method does not represent or imply that all steps must be executed in this order. A person skilled in the art can transform or change the execution order of the above steps based on the present invention. Some embodiments of the above method are illustrated below.
[0174] See also Figure 7 This embodiment also provides a system for designing and evaluating whole virus primers, the system comprising:
[0175] An acquisition module, used to obtain a viral genome sequence library of a virus;
[0176] The first design evaluation module is used to design usable primers for species identification using Scheme 1 or Scheme 2, perform species verification and host verification on the usable primers for species identification, and evaluate each pair of usable primers for species identification;
[0177] The second design evaluation module is used to design usable primers for subtype identification using Scheme 1 or Scheme 2, perform non-target virus subtype verification, interspecies verification, and host verification on the usable primers for subtype identification, and evaluate each pair of usable primers for subtype identification;
[0178] The first scheme is: determining a representative genome of the virus in a viral genome sequence library, designing primers for the representative genome, using a non-representative genome in the viral genome sequence library to specifically verify the primers for the representative genome, and using the primers for the representative genome that meet the verification pass rate as available primers;
[0179] The second scheme is: perform multiple sequence alignment on the genome sequences in the viral genome sequence library to identify the conserved regions of the virus, design primers for the conserved regions, use the viral genome sequence library to perform specific verification on the primers in the conserved regions, and use the primers in the conserved regions with a verification pass rate that meets the standard as available primers.
[0180] The first design evaluation module includes: a first design unit, a species identification verification unit, and a first evaluation unit. The first design unit is used to design available primers for species identification using scheme 1 or scheme 2, the species identification verification unit is used to perform species verification and host verification on the available primers for species identification, and the first evaluation unit is used to evaluate each pair of available primers for species identification.
[0181] The second design evaluation module includes: a second design unit, a subtype identification verification unit, and a second evaluation unit. The second design unit is used to design available primers for subtype identification using scheme 1 or scheme 2, the subtype identification verification unit is used to perform inter-subtype verification and host verification on the available primers for subtype identification, and the second evaluation unit is used to evaluate each pair of available primers for subtype identification.
[0182] When the system for designing and evaluating whole-virus primers is implemented, the design and evaluation of primers can be implemented by referring to the method for designing and evaluating whole-virus primers in any of the above embodiments, and the specific implementation steps will not be repeated here.
[0183] The present invention also provides a storage medium storing a computer program, wherein the storage medium includes instructions, and when the computer program is executed by a processor, the steps of the method for designing and evaluating a whole-virus primer are implemented.
[0184] The effects of the method, system and storage medium for designing and evaluating a whole virus primer of the present invention are as follows: the present invention is a method and system for designing and evaluating a joint primer for virus species identification and a primer for virus subtype identification, and corresponding primers can be designed for different subtypes of the virus, respectively, and the design of primers for virus subtype identification is universal and highly accurate. The present invention obtains usable primers by designing primers and verifying the specificity of the designed primers, and the primers are designed based on conserved regions, which reduces the probability of amplification failure, and the usable primers for species identification and subtype identification are fully verified, and analysis is given based on the perspectives of species, species, (subtype) and host, and the primers are evaluated. Such a design makes the present invention suitable for the design of virus primers with different application scenarios and different genomic characteristics, and on the scale of the whole virus, the high specificity of the candidate primers for the virus is guaranteed (to the maximum extent), and the accuracy of the primers is ensured (to the maximum extent). The primers of the present invention are designed based on conserved regions, which reduces the probability of amplification failure. The present invention comprehensively considers the perspectives of intra-species, inter-species, (inter-subtype) and host, and the reference information is more complete, which improves the verification dimension and evaluation accuracy, thereby reducing the application risk; provides the application idea of primers across species; provides specific verification in different hosts, and based on the present invention, primers can be selected according to the host, reducing the application risk. The whole virus primer set for all viral genomes within the species and subtype range that can be formed according to this scheme (maximum) ensures the accuracy of the primers and the breadth of the detectable virus spectrum.
[0185] A specific embodiment of the method of the present invention is given below.
[0186] The implementation process is described using Semliki Forest virus (taxid 11033) as an example.
[0187] (1) Genome sequence screening.
[0188] A total of 4 genome sequences of the virus's GenBank Assembly and RefSeq Assembly were obtained from NCBI, and the corresponding Accession IDs were: GCA_000860285.1, GCA_002889475.1, GCA_031161755.1 and GCF_000860285.1. Among them, GCF_000860285.1 and GCF_000860285.1 are the same genome. Therefore, the three genomes of GCF_000860285.1, GCA_002889475.1, and GCA_031161755.1 were used to enter the next module.
[0189] (2) Primer design for species identification
[0190] The virus has a RefSeq genome sequence, and primers for species identification were designed.
[0191] The primers were designed by designating the genome GCF_000860285.1 as the representative genome. GCA_002889475.1 and GCA_031161755.1 were the neighbor genomes. Table 3 shows the grouping of the Semliki Forest virus genomes.
[0192] Table 3
[0193] GCF_000860285.1 is 11442 bp in length and was cut into two sequences of 10000 bp in length at a step length of 9500 bp to obtain two sequences, NC_003215_001_1_10000.fa and NC_003215_002_9501_11442.fa.
[0194] Primers were designed for these two sequences using primer3 software, with the following main parameters: primer length 18bp-23bp, product length 100bp-300bp, RCR annealing temperature 59℃, GC content range 30%-70%. The number of primers returned is the integer obtained by rounding off the value of "template sequence length (bp) / 50", i.e., NC_003215_001_1_10000.fa setting returns 100 pairs of primers, and NC_003215_002_9501_11442.fa setting returns 39 pairs of primers.
[0195] Since the two sequences have a 500 bp duplication region, 7 pairs of identical primers in this region were predicted by both sequences at the same time. After deduplication, a total of 232 pairs of primers were obtained. The MFEprimer software (default parameters) was used to perform a specificity test on the whole genome of GCF_000860285.1, and the results showed that all primers had amplification products and the target was unique.
[0196] Based on the genome sequence and the results of primer3, the information of 232 pairs of primers was extracted. The extracted content and field description are shown in Table 4, which shows the primer information of the representative genome of the virus.
[0197] The MEFPrimier software (default parameters) was used to test the specificity of the 232 primer pairs obtained on each neighbor's genome, and the blast software (evalue = 1e-10) was used to align the product sequences. The internal code was used to analyze the validation of each primer pair on each neighbor's genome. Taking GCA_002889475.1 as an example, some examples of primer validation results are shown in Table 5.
[0198] Table 4
[0199] Table 5
[0200] Table 6
[0201] The passing rate of each pair of primers on the neighbors genome was calculated, and the passing rates are shown in Table 6.
[0202] Primers with a passing rate of no less than 50% for the neighbors genome were screened from Table 6, totaling 230 pairs, which were used as species primers for subsequent analysis.
[0203] Construct a representative viral sequence database. Use MEFPrimer software with default parameters, input the above 230 primer sets and the representative viral sequence database, and output the situation in which the primers can amplify products in each genome of the virus library. For each pair of primers, count the situations in which products can be amplified on sequences of non-species. The results of the inter-species validation of the virus library are shown in Table 7 below.
[0204] Table 7
[0205] In Table 7, the primer_id column is the name of the primer pair, the target_virus_taxon_num is the number of non-species genomes on which the primer can amplify, and the target_virus_info column is a specific description of the non-species, with species separated by commas “,”. The information of each species, with “|” as a connector, represents the taxid, Organism_name, Species_name, and the number of products (hit_num) that the primer can amplify on the species.
[0206] Host verification: The main hosts of the virus in vertebrates are humans and sheep. For these two hosts, the MEFPrimer software was used with default parameters, and the primer set and host genome were input to test whether the primers can amplify products on the host genome. The host genome verification results are shown in Table 8 below.
[0207] In Table 8, t_h_n represents target_host_number, which is the number of hosts on which the primer can amplify; the target_host_info column is a detailed description of the host, and the numbers in brackets represent the number of amplification targets of the primer on the host genome.
[0208] Table 8
[0209] Primer evaluation was performed, and some of the evaluation results are shown in Table 9.
[0210] Table 9
[0211] Among them, passed_neighbors_number indicates the number of neighbors genomes that have passed, and Level indicates the rating result.
[0212] The statistics of primer numbers at each level are shown in Table 10:
[0213] Table 10
[0214] The examples of primers for subtype identification of Semliki Forest virus obtainable according to the above method are not described here.
[0215] The following are primers for identifying Norovirus H subtypes, and the specificity of the primers for other subtypes needs to be considered (this example uses four subtypes, I, F, J, and G, as examples for analysis). The process and results of the primers for identifying the species of Norovirus are not shown here.
[0216] For norovirus genome sequence screening, see Table 11.
[0217] Table 11
[0218] This embodiment is to design primers for subtypes of Norovirus. According to prior knowledge, the viral genome is composed of 11 segments, and the difference between different subtypes is mainly the difference in the VP6 gene. In order to enable the primers to avoid interference from other subtype genomes, the VP6 gene of the target subtype is directly extracted, and then the primers are designed. The target subtype Rotavirus_H has only one genome and does not involve the neighbors genome. The primers designed based on the VP6 gene are specifically verified by the whole genome of Rotavirus_H, and the only primers at the amplification position are retained, totaling 26 pairs of primers. Examples of primers for Rotavirus_H subtype identification are shown in Table 12.
[0219] Non-target subtype genome verification: Using MEFPrimer software and default parameters, the 26 pairs of primers obtained above were verified on the other 4 non-target subtypes, a total of 5 genomes. The number of amplified targets and subtype information of each pair of primers on the non-target subtype genome were used for subsequent primer evaluation. In this example, none of the 26 pairs of primers had an amplification effect on the non-target subtype genome.
[0220] Validation of virus library species: Use MEFPrimer software with default parameters, input the above 26 primer sets and representative virus sequence database, and output the situation that the primers can amplify products in each genome of the virus library. For each pair of primers, the situation that the product can be amplified on the sequence of non-species is counted. The results of the virus library species validation of Rotavirus_H identification primers are shown in Table 13.
[0221] Host verification: The main host of norovirus is human. In this embodiment, primer verification was performed on the human genome. MEFPrimer software was used with default parameters, and the 26 pairs of primers obtained above and the representative viral sequence database were input to test whether the primers could amplify products on the host genome. The results are shown in Table 14 below.
[0222] Table 12
[0223] Table 13
[0224] Table 14
[0225] Primer evaluation, the results of the Rotavirus_H subtype identification primers are shown in Table 15.
[0226] Table 15
[0227] Among them, neighbors_nubmer indicates the number of neighbors genomes, and target_other_subtype_num indicates how many non-target subtype genomes the primer can amplify.
[0228] The number of primers at each level was counted, and the results of primer evaluation for Rotavirus_H subtype identification primers are shown in Table 16.
[0229] Table 16
[0230] It should be noted that in the above embodiments, the description of each embodiment has its own emphasis, and for parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0231] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or storage media. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0232] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system and method according to multiple embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a part of a module, a program segment or a code, and the part of the module, the program segment or the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous square boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or the flow chart, and the combination of the square boxes in the block diagram and / or the flow chart can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0233] Although preferred embodiments of the present invention have been described, additional changes and modifications may be made to these embodiments by those skilled in the art once the basic inventive concept is known. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention. Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A method for designing and evaluating whole virus primers, characterized in that: include: obtaining a viral genome sequence library of the virus; Designing usable primers for species identification and / or designing usable primers for subtype identification; The method of designing the available primers for species identification is as follows: using scheme 1 or scheme 2 to design the available primers for species identification, performing species verification and host verification on the available primers for species identification, and evaluating each pair of available primers for species identification; The method of designing available primers for subtype identification is as follows: using scheme 1 or scheme 2 to design available primers for subtype identification, performing non-target virus subtype verification, species verification and host verification on the available primers for subtype identification, and evaluating each pair of available primers for subtype identification; The first scheme is: determining a representative genome of the virus in a viral genome sequence library, designing primers for the representative genome, using a non-representative genome in the viral genome sequence library to specifically verify the primers for the representative genome, and using the primers for the representative genome that meet the verification pass rate as available primers; The second scheme is: perform multiple sequence alignment on the genome sequences in the viral genome sequence library to identify the conserved regions of the virus, design primers for the conserved regions, use the viral genome sequence library to perform specific verification on the primers in the conserved regions, and use the primers in the conserved regions with a verification pass rate that meets the standard as available primers.
2. The method for designing and evaluating a whole virus primer according to claim 1, characterized in that: The viral genome sequence library is a processed viral genome sequence library, and the viral genome sequence library of the virus is obtained specifically by: obtaining the virus, collecting the subtype information, genome sequence and genome sequence metadata of the virus genome to form an original viral genome sequence library, and performing sequence deduplication and filtering processing on the original viral genome sequence library to obtain the processed viral genome sequence library.
3. The method for designing and evaluating a whole virus primer according to claim 1, characterized in that: If the virus is a single genome virus, use Scheme 1 to design primers that can be used for species identification and primers that can be used for subtype identification; If the virus is a non-single-genome virus, the available primers for species identification designed using Scheme 1 or Scheme 2 are determined based on the variability of the non-single-genome virus and / or the number of genomes of the non-single-genome virus, and the available primers for subtype identification designed using Scheme 1 or Scheme 2 are determined based on the variability of the non-single-genome virus and / or the number of genomes of the non-single-genome virus.
4. The method for designing and evaluating a whole virus primer according to claim 1, characterized in that: The specificity verification specifically includes: using the verification genome to perform specificity analysis on the primers representing the genome, determining whether the primers representing the genome meet the specificity standard to obtain a specificity verification result, and calculating the specificity verification pass rate of the primers representing the genome according to the specificity verification result; In the scheme 1, the verification genome is a non-representative genome in the viral genome sequence library. If the virus does not have a non-representative genome, the primers representing the genome are deemed to have passed specific verification, and the verification pass rate is 100%. If the virus has a non-representative genome, the non-representative genome in the viral genome sequence library is used as the verification genome; In the second scheme, the verified genome is the genome of the viral genome sequence library.
5. The method for designing and evaluating a whole virus primer according to claim 4, characterized in that: The specificity standard is to meet all of the following conditions: Condition 1: The primers can have an amplification effect on the verified genome and the amplification position is unique; Condition 2: The gene amplified by the primers on the verification genome is consistent with the gene amplified by the primers on the representative genome. If the verification genome is a segment genome, the amplification product of the primers on the verification genome and the amplification product of the primers on the verification genome are located in the same segment; Condition 3: Verify that the length of the amplified product on the genome and the length variation of the amplified product of the primer on the representative genome are within the set length variation threshold range; Condition 4: Verify that the sequence similarity between the amplified product on the genome and the amplified product of the primer on the representative genome reaches the set sequence similarity threshold condition.
6. The method for designing and evaluating a whole virus primer according to claim 1, characterized in that: The primers available for species identification for virus species verification and host verification include: Amplification of the available primers for species identification in the sequences of the representative virus sequence database, and statistical amplification to obtain the virus species verification results of the available primers for species identification, wherein the representative virus sequence database includes representative sequences of multiple virus species, Primers available for species identification are amplified on all host genomes in a host genome database, and the amplification conditions are statistically analyzed to obtain host verification results of the primers available for species identification, wherein the host genome database includes genomes of several hosts of the virus.
7. The method for designing and evaluating a whole virus primer according to claim 1, characterized in that: The subtype identification primers used for non-target virus subtype verification, virus species verification and host verification specifically include: The available primers for subtype identification are amplified on the genome of the non-target subtype, and the amplification situation is statistically analyzed to obtain the verification result of the non-target virus subtype; The available primers for subtype identification are amplified on all representative sequences of a representative viral sequence database, and the amplification status is statistically analyzed to obtain the viral species verification results of the available primers for subtype identification, wherein the representative viral sequence database includes representative sequences of multiple viral species; The available primers for subtype identification are amplified on all host genomes in the host genome database, and the amplification conditions are statistically analyzed to obtain host verification results of the available primers for subtype identification. The host genome database includes genomes of several hosts of the virus.
8. The method for designing and evaluating a whole virus primer according to claim 1, characterized in that: The evaluation of each pair of available primers for species identification is specifically as follows: for each pair of available primers for species identification, quality rating is performed based on its verification pass rate, its virus species verification result and its host verification result; The evaluation of each pair of available primers for subtype identification is specifically as follows: for each pair of available primers for subtype identification, quality rating is performed based on its verification pass rate, its non-target virus subtype verification result, its virus species verification result and its host verification result.
9. A design and evaluation system for whole virus primers, characterized in that: include: The first design evaluation module is used to design usable primers for species identification using Scheme 1 or Scheme 2, perform species verification and host verification on the usable primers for species identification, and evaluate each pair of usable primers for species identification; The second design evaluation module is used to design usable primers for subtype identification using Scheme 1 or Scheme 2, perform non-target virus subtype verification, interspecies verification, and host verification on the usable primers for subtype identification, and evaluate each pair of usable primers for subtype identification; The first scheme is: determining a representative genome of the virus in a viral genome sequence library, designing primers for the representative genome, using a non-representative genome in the viral genome sequence library to specifically verify the primers for the representative genome, and using the primers for the representative genome that meet the verification pass rate as available primers; The second scheme is: perform multiple sequence alignment on the genome sequences in the viral genome sequence library to identify the conserved regions of the virus, design primers for the conserved regions, use the viral genome sequence library to perform specific verification on the primers in the conserved regions, and use the primers in the conserved regions with a verification pass rate that meets the standard as available primers.
10. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method for designing and evaluating a whole virus primer as described in any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Design method and system of targeted pathogenic microorganism sequencing primer
CN118762752A
Oligonucleotide probes for specific identification of noroviruses and other pathogens
US20160034636A1
Method and apparatus for determining microbial species and acquiring related information by means of sequencing, computer-readable storage medium, and electronic device
WO2022028624A1