Target virus primer generation method and device, medium and product
By obtaining the viral genome and generating primer design index values based on primer design patterns, the problem of primer design relying on manual methods has been solved, realizing the automation and precision of viral primer design, and adapting target viral primers to the viral genome, thereby improving the efficiency and accuracy of primer design.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BERGER (QINGDAO) MEDICAL TECH CO LTD
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-21
AI Technical Summary
Existing primer design software cannot achieve automatic sequence quality control and design region selection, resulting in time-consuming and labor-intensive primer design and problems such as insufficient coverage or specificity.
By acquiring the genome of the virus to be analyzed, analysis is performed based on primer design patterns to generate primer design index values. Target viral primers adapted to the viral genome are designed automatically, including identifying conserved regions and compatibility with viral subtype variations. Software tools such as seqkit, cd-hit, MAFFT, and blastn are used for sequence processing and alignment.
It achieves automation, efficiency, and precision in viral primer design, ensuring primer specificity and coverage, and enabling accurate identification and detection of different viral subtypes.
Smart Images

Figure CN121905263A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of molecular biology, and more particularly to a method, apparatus, medium, and product for generating target virus primers. Background Technology
[0002] Viral genomes exhibit high mutation rates, with significant differences even among different subtypes within the same viral species. Take enteroviruses as an example: they are classified into four species (A, B, C, and D), and further subdivided into Coxsackievirus, Echovirus, Poliovirus, and others, totaling over 100 subtypes. This presents a significant challenge to primer design; how to screen for conserved regions and design primer combinations that cover a wider range of types is a major hurdle.
[0003] Currently, commonly used primer design software includes Primer3 and Oligo7. These software programs can output candidate primers by setting parameters such as primer size, amplification product size, Tm value, and GC value, and can also perform preliminary assessments of the secondary structure risks of candidate primers, such as hairpin structures and primer dimers.
[0004] However, currently used primer design software cannot complete the tasks of sequence quality control and automatic selection of design regions. These two core tasks still need to be completed manually, which is not only time-consuming and labor-intensive, but also prone to insufficient primer coverage or specificity due to differences in human judgment. Summary of the Invention
[0005] This disclosure provides a method, apparatus, medium, and product for generating target virus primers, which solves the problem of relying on manual methods for sequence quality control and design region selection, and realizes the automation, efficiency, and precision of virus primer design.
[0006] According to one aspect of this disclosure, a method for generating target virus primers is provided, comprising:
[0007] Obtain the genome of the virus to be analyzed;
[0008] Based on the primer design pattern of the viral genome to be analyzed, the viral genome to be analyzed is analyzed to obtain primer design index values;
[0009] Based on the primer design index values, target viral primers adapted to the genome of the virus to be analyzed are generated.
[0010] According to another aspect of this disclosure, a target virus primer generation apparatus is provided, the apparatus comprising:
[0011] The module for acquiring the genome of the virus to be analyzed is used to acquire the genome of the virus to be analyzed.
[0012] The primer design index value acquisition module is used to analyze the genome of the virus to be analyzed based on the primer design pattern of the genome of the virus to be analyzed, and obtain the primer design index value.
[0013] The target virus primer acquisition module is used to generate target virus primers adapted to the genome of the virus to be analyzed based on the primer design index values.
[0014] According to another aspect of this disclosure, an electronic device is provided, the electronic device comprising:
[0015] At least one processor; and
[0016] A memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the target virus primer generation method according to any embodiment of this disclosure.
[0018] According to another aspect of this disclosure, a computer-readable storage medium is provided that stores computer instructions for causing a processor to execute and implement the target virus primer generation method according to any embodiment of this disclosure.
[0019] According to another aspect of this disclosure, a computer program product is provided, which, when executed by a processor, implements the target virus primer generation method as described in any of the embodiments of this disclosure.
[0020] This embodiment of the disclosure obtains the genome of a virus to be analyzed; analyzes the genome based on the primer design pattern of the genome to be analyzed to obtain primer design index values; and generates target viral primers adapted to the genome to be analyzed based on the primer design index values. By analyzing the viral genome and obtaining primer design index values, this invention solves the problem of relying on manual methods for sequence quality control and design region selection, and realizes the automation, efficiency, and precision of viral primer design.
[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0022] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of this disclosure and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a flowchart of a method for generating target virus primers according to an embodiment of this disclosure;
[0024] Figure 2 This is a flowchart of a method for generating target virus primers according to an embodiment of this disclosure;
[0025] Figure 3 This is a flowchart of the virus typing primer design pattern in the embodiments of this disclosure;
[0026] Figure 4 This is a flowchart of the viral whole genome primer design in the embodiments of this disclosure;
[0027] Figure 5 This is a schematic diagram of the structure of a target virus primer generation device according to an embodiment of this disclosure;
[0028] Figure 6 This is a schematic diagram of the structure of an electronic device according to an embodiment of this disclosure. Detailed Implementation
[0029] To enable those skilled in the art to better understand the present disclosure, the technical solutions of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present disclosure, and not all embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present disclosure.
[0030] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0031] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0032] Figure 1 This flowchart illustrates a method for generating target virus primers according to an embodiment of the present disclosure. This embodiment is applicable to the design of genotyping primers and multiple primers targeting viral genomes. The method can be executed by the target virus primer generation device described in this disclosure. This device can be implemented using software and / or hardware and can be integrated into electronic devices such as computer equipment, servers, mobile terminals, or processors. Figure 1 As shown, the method specifically includes the following steps:
[0033] S110, Obtain the genome of the virus to be analyzed.
[0034] In this embodiment, the viral genome to be analyzed can be understood as the set of nucleic acid sequences of the target virus for which primers need to be designed. It can be a known viral reference genome sequence downloaded from an open-source database. The core purpose is to screen conserved regions, calculate primer design indicators, and generate target primers through subsequent sequence analysis. The open-source database can be NCBI (National Center for Biotechnology Information).
[0035] Specifically, download the genome of the virus to be analyzed from an open-source database, where primers need to be designed, to ensure that subsequent analysis focuses on the core target and avoids interference from irrelevant sequences.
[0036] Optionally, the genomic sequences in the viral genome to be analyzed are subjected to quality testing, and genomic sequences that do not meet the quality conditions are removed. In this embodiment, genomic sequences that do not meet the quality conditions can be specifically understood as sequences whose N base ratio exceeds a preset threshold, wherein the preset threshold can be 20%.
[0037] Specifically, the genome sequences in the viral genome to be analyzed are subjected to quality testing, and sequences with redundant or N base ratios exceeding a preset threshold are filtered out and sorted from high to low according to sequence length.
[0038] S120, Based on the primer design pattern of the viral genome to be analyzed, the viral genome to be analyzed is analyzed to obtain primer design index values.
[0039] In this embodiment, primer design mode can be specifically understood as a specific design mode adapted to different primer design goals. Primer design mode can include, but is not limited to, virus typing primer design mode, used to guide the automated processing of the viral genome to be analyzed and the generation of primer design indicators. Primer design indicator values can be specifically understood as multi-dimensional indicators supporting the generation of target primers. Primer design indicator values can include, but are not limited to, primer length, amplification product length, Tm (melting temperature) value, etc., including both the basic physicochemical parameters of the primer itself and the core performance parameters such as specificity and coverage of different design goals. They can be directly used as the screening basis for qualified primers, thereby ensuring the practical effectiveness and adaptability of the final primers.
[0040] Specifically, based on the primer design pattern of the viral genome to be analyzed, the viral genome is analyzed to obtain primer design index values. By matching the primer design pattern with the characteristics of the viral genome to be analyzed, targeted analysis is performed to ensure that the primer design index values are accurately adapted to the viral genome structure and requirements, reducing blind screening, improving the specificity, coverage and experimental success rate of target primers, and providing a standardized and personalized technical path.
[0041] Optionally, based on the above embodiments, the viral genome to be analyzed includes multiple subtype genomes, and the primer design mode includes a viral typing primer design mode, which is used to generate target viral primers corresponding to each viral subtype.
[0042] In this embodiment, the viral genome to be analyzed includes multiple subtype genome sequence sets. Multiple subtype genome sequence sets refer to multiple genomes corresponding to different types within the same viral species. Each subtype genome sequence set represents the sequence characteristics of that type and contains a sufficient number of sequences to cover the internal variations of that subtype. Primer design patterns include viral typing primer design patterns. Viral typing primer design patterns refer to design patterns specifically for each subtype within the same viral species, with the core objective of generating primers for each subtype that bind specifically to its own subtype genome and are compatible with the variations of strains within that subtype, thereby achieving accurate differentiation between different viral types. For example, CA10 (Coxsackievirus A10), CA16 (Coxsackievirus A16), and EV71 (Enterovirus A71) all belong to group A of enteroviruses, but their severity risk levels differ. For primer design of these three subtypes, viral typing primer design patterns can be used.
[0043] Optionally, based on the above embodiments, the genome of the virus to be analyzed is analyzed to obtain primer design index values, including: identifying the representative genome sequence of each viral subtype; segmenting the genome sequences in the multiple viral subtype genome sequence set to obtain multiple sequence segments; comparing the multiple sequence segments with the representative genome sequence of each viral subtype to obtain alignment results; determining conserved regions in the genome of each viral subtype based on the alignment results; and determining the primer design index values corresponding to each viral subtype genome based on the conserved regions in the genome of each viral subtype.
[0044] In this embodiment, the representative genome sequence can be understood as the sequence in each viral subtype genome sequence set that most accurately reflects the core sequence characteristics of that subtype and is as compatible as possible with variations of different strains within that subtype. The representative genome sequence must possess high sequence integrity, no missing core functional genes, and strong coverage of variant sites. It can serve as a representative of the viral subtype genome for subsequent sequence segmentation and conserved region screening. A sequence segment can be understood as a sequence fragment formed by splitting each genome sequence from multiple viral subtype genome sequence sets into a preset fixed length. By setting a reasonable step size, conserved region breaks are avoided, providing a standardized analysis unit for subsequent accurate alignment with the representative genome sequence and conserved region localization. For example, the preset fixed length can be 300 bp.
[0045] The alignment results can be understood as the results obtained by comparing multiple sequence segments with the representative genome sequence of the corresponding viral subtype. The core purpose is to reflect the similarities and differences between each sequence segment and the representative genome sequence, providing a direct basis for subsequent screening of conserved regions. Conserved regions can be specifically understood as the base sequences selected from the genome sequence set of each viral subtype after comparing multiple sequence segments with the representative genome sequence. Their core characteristics are universality within the subtype, strong sequence stability, and the ability to meet the design requirements of conserved primer sets, providing core target regions for subsequent design of specific primers adapted to all strains of that subtype.
[0046] Specifically, the representative genomic sequence of each viral subtype is identified, and the genomic sequences from multiple viral subtype genomic sequence sets are cut into short fragments of the same predetermined fixed length, resulting in multiple sequence segments. The seqkit tool can be used for this cutting. These multiple sequence segments are then aligned with the representative genomic sequence of each viral subtype to obtain alignment results. Based on the alignment results, conserved regions in each viral subtype are identified, and primer design metrics for each subtype's genome are determined based on these conserved regions. The entire process is automated through standardized splitting, alignment, and screening logic, eliminating the need for manual intervention in region selection and balancing the efficiency and reliability of primer design.
[0047] Optionally, based on the above embodiments, identifying the representative genome sequence of each viral subtype includes: performing clustering processing on the genome sequence set of each subtype to obtain clustering results, wherein the clustering results include at least one sequence cluster set, and each sequence cluster set includes at least one genome sequence; determining the sequence cluster set with the largest number of genome sequences, and determining the longest genome sequence in the sequence cluster set with the largest number of genome sequences as the representative genome sequence of the viral subtype.
[0048] In this embodiment, the clustering result can be specifically understood as the classification result obtained after clustering all genomic sequences in the same subtype genomic sequence set. Genomic sequences with a sequence similarity exceeding a preset value are grouped into the same sequence cluster set. For example, the preset value can be 0.9. Genomic sequences with a sequence similarity exceeding 0.9 are grouped into the same sequence cluster to form multiple non-overlapping sequence cluster sets, which are used to screen the most representative core sequence clusters in this viral subtype. The sequence cluster set can be specifically understood as the set of genomic sequences in the same subtype genome that reach the preset value after clustering, and the sequences in different sequence cluster sets have significant differences.
[0049] Specifically, the genomic sequences in each viral subtype's genomic sequence set are clustered using CD-HIT clustering software. From these clustering results, the cluster with the largest number of genomic sequences is identified, and the longest genomic sequence within this cluster is selected as the representative genomic sequence for the viral subtype. Choosing the longest genomic sequence from this cluster ensures the integrity of the representative genomic sequence and avoids biases in subsequent alignment and conserved region selection due to sequence fragmentation.
[0050] Optionally, based on the above embodiments, the alignment result includes the positional depth of the sequence segment.
[0051] Specifically, the depth of the locus can be understood as, after comparing multiple sequence segments with the representative genome sequence of the corresponding viral subtype, counting the number of sequence segments that match each base site on the representative genome sequence. The core purpose is to quantify the coverage and prevalence of each base site within the genotyping, providing supplementary judgment criteria for the accurate screening of conserved regions.
[0052] Optionally, based on the above embodiments, determining the conserved region of each viral subtype based on the alignment results includes: determining the intragroup depth of each sequence segment in the representative genome sequence of its viral subtype and the outgroup depth of each sequence segment in the representative genome sequences of other viral subtypes based on the alignment results; filtering the intragroup depth of each sequence segment based on an intragroup depth threshold and filtering the outgroup depth of each sequence segment based on an outgroup depth threshold to obtain continuous sequence segments that satisfy the intragroup depth threshold and the outgroup depth threshold, and determining the continuous sequence segments as the conserved region.
[0053] In this embodiment, intragroup depth can be specifically understood as the total number of sequence segments that match the base sites of the representative genome sequence covered by the sequence segment after alignment with the representative genome sequence of the viral subtype. This is crucial for quantifying the coverage and conservation of the sequence segment within its viral subtype, directly reflecting whether the sequence segment is highly homologous within the same subtype, and is a key quantitative indicator for determining intragroup conservation. Outgroup depth can be specifically understood as the total number of heterogenous sequence segments that match the base sites of the representative genome sequences of other subtypes covered by the sequence segment after alignment with the representative genome sequence of other subtypes outside its subtype. This is crucial for quantifying the cross-homology of the sequence segment with other subtypes, and is a key quantitative indicator for determining intergroup specificity.
[0054] The intragroup depth threshold can be understood as a preset depth threshold used to determine whether a sequence segment meets the intragroup conservation requirement. It is used to screen sequence segments with high coverage and broad applicability within the same subtype, ensuring that subsequent conserved regions meet the primer adaptation requirements of the same subtype. The extragroup depth threshold can be understood as a preset depth threshold used to determine whether a sequence segment meets the intergroup specificity requirement. It is used to screen sequence segments with low cross-homology with other subtypes and high specificity, preventing non-specific binding of subsequently designed primers to other viral subtypes. For example, if the intragroup depth threshold is 40 and the extragroup depth threshold is 3, then a sequence segment must simultaneously meet the requirement of an intragroup depth greater than or equal to 40 and an extragroup depth less than 3 to be identified as a conserved region.
[0055] Specifically, based on the SAM format alignment file, the intra-group depth of each sequence segment within the representative genome sequence of its subtype and the out-of-group depth of each sequence segment within the representative genome sequences of other subtypes are determined. Intra-group depth thresholds are used to filter the intra-group depth of each sequence segment, and out-of-group depth thresholds are used to filter the out-of-group depth of each sequence segment, resulting in conserved regions that meet both intra-group and out-of-group depth thresholds. Quantifying depth ensures the core characteristics of conserved regions: "intra-group conservation and inter-group specificity."
[0056] Optionally, based on the above embodiments, the alignment results also include a base distribution matrix.
[0057] In this embodiment, the base distribution matrix can be specifically understood as a matrix that presents the base type and frequency of each alignment site after comparing multiple sequence segments with the corresponding representative genome sequence. The rows of the matrix correspond to the coordinates of the base sites of the representative genome sequence, and the columns correspond to the base types. This matrix is used to intuitively reflect the base conservation and variation distribution characteristics of each site, providing a refined basis at the base level for screening conserved regions.
[0058] Optionally, based on the above embodiments, the intragroup consistency verification of the continuous sequence segments in the genome sequence set of their respective viral subtypes is performed based on the base distribution matrix.
[0059] Specifically, the intra-group consistency of continuous sequence segments within the genome sequence set of their respective viral subtypes is verified based on the base distribution matrix. By analyzing the proportion of bases at each point in the matrix, highly conserved sites with a high proportion of a single base and variant sites with a balanced distribution of multiple bases are quickly identified, avoiding the neglect of base mismatch issues by only looking at intra-group depth.
[0060] Optionally, based on the above embodiments, the primer design index value corresponding to each subtype genome is determined based on the conserved region in each viral subtype, including: extracting the genome sequence segments relative to the conserved region from the genome sequence set of each viral subtype to form a sequence set; performing multiple sequence alignment on multiple sequence segments in the sequence set to obtain multiple sequence alignment results; and determining the primer design index value corresponding to each viral subtype based on the multiple sequence alignment results.
[0061] In this embodiment, the sequence set can be specifically understood as a collection of sequence fragments extracted from all genomic sequences of the same subtype that completely correspond to the identified conserved regions of that subtype. All sequence segments within the set precisely match the coordinate range of the conserved regions, providing standardized input data for subsequent multiple sequence alignment and primer design index calculations. The multiple sequence alignment result can be specifically understood as the sequence alignment result obtained after aligning all sequence segments in the sequence set. The sequence segments in the set are arranged according to their base sites, visually presenting the consistency and distribution of variant sites at each base site within the conserved regions, providing a precise sequence alignment basis for the quantitative calculation of primer design index values.
[0062] Specifically, the BLASTN alignment tool can be used to extract sequence segments from the genome sequence set of each viral subtype relative to conserved regions, forming a sequence set. Multiple sequence alignments are then performed on multiple sequence segments extracted from the sequence set of each viral subtype using the MAFFT multiple sequence alignment tool to obtain the alignment results. Based on the multiple sequence alignment results, primer design index values corresponding to each viral subtype are determined. The primer design index values are quantitatively derived based on the alignment results, replacing empirical design and ensuring a high degree of matching between the index values and the actual physicochemical properties and variation patterns of the subtyped sequences.
[0063] It should be noted that when using the MAFFT multiple sequence alignment tool to obtain multiple sequence alignment results, a corresponding degenerate base can be set according to the degenerate base introduction threshold. For example, when the proportion of A bases at a certain site is 40%, T bases 35%, C bases 20%, and G bases 5%, and the preset degenerate base introduction threshold is "containing more than 70%", a degenerate base W (corresponding to both A and T bases) can be introduced at that site. Introducing degenerate bases based on the base proportion threshold can avoid primer failure due to base variations at individual sites.
[0064] S130, Based on the primer design index values, generate target viral primers adapted to the genome of the virus to be analyzed.
[0065] In this embodiment, the target virus primers can be specifically understood as primers designed and generated based on primer design index values. These primers target conserved regions of the genome of each subtype of the virus to be analyzed, are precisely adapted to the target subtype, are compatible with intra-subtype variations, and have no cross-subtype reactions. They can be used as core tools for experiments such as PCR (Polymerase Chain Reaction) amplification and nucleic acid detection, enabling accurate identification and detection of the target virus subtype.
[0066] Specifically, target viral primers adapted to the genome of the virus to be analyzed are generated based on the primer design index values corresponding to each viral subtype. Generating target viral primers based on the primer design index values corresponding to the viral subtype can ensure the subtype specificity and amplification effectiveness of the primers, and achieve accurate differentiation of different viral subtypes.
[0067] This embodiment of the disclosure obtains the genome of a virus to be analyzed; analyzes the genome based on the primer design pattern of the genome to be analyzed to obtain primer design index values; generates target viral primers adapted to the genome to be analyzed based on the primer design index values, and analyzes the viral genome through the primer design pattern to obtain primer design index values. This solves the problem of relying on manual methods for sequence quality control and design region selection, and realizes the automation, efficiency and accuracy of viral primer design.
[0068] Figure 2 This is a flowchart illustrating a method for generating target virus primers according to an embodiment of the present disclosure. This embodiment provides a flowchart of a viral whole-genome primer design mode within the target virus primer generation method. The viral genome to be analyzed in the viral whole-genome primer design mode includes the entire viral genome. The viral whole-genome primer design mode is used to generate target virus primers corresponding to a specified type of virus, adapting to the multiple primer design requirements of the viral whole genome. For example... Figure 2 As shown, the method specifically includes the following steps:
[0069] S210, Obtain the genome of the virus to be analyzed.
[0070] Optionally, if it is necessary to process multiple subtypes contained in the viral genome sequence to be analyzed separately, an average base similarity threshold can be set to cluster the multiple genome sequences of the virus to be analyzed and divide them into multiple sequence sets for primer design.
[0071] S220, perform sequence alignment on the genome of the virus to be analyzed to obtain a consistent sequence of the genome of the virus to be analyzed.
[0072] In this embodiment, the homogeneous sequence of the viral genome to be analyzed can be specifically understood as the genome sequence containing degenerate bases obtained from the complete genome sequence of the virus to be analyzed through sequence alignment. This sequence can accurately reflect the main characteristics of the virus. Its core function is to participate in subsequent sliding window traversal, sequence segment analysis, and primer design index calculation, which reduces computational complexity and accurately represents the overall sequence characteristics of the virus.
[0073] Specifically, sequences in the viral genome to be analyzed are compared to obtain a consistent sequence. This provides an accurate and representative reference template for subsequent primer design, avoiding primer design bias caused by an excessive number of sequences or dispersed features, while ensuring primer coverage and amplification specificity.
[0074] Optionally, if the length of the genome sequence of the virus to be analyzed is less than or equal to a preset length threshold and the number of genome sequences is less than or equal to a preset number threshold, MAFFT software can be used to obtain a consistent sequence when comparing the sequences in the genome of the virus to be analyzed.
[0075] If the length of the genome sequence of the virus to be analyzed exceeds a preset length threshold and / or the number of genome sequences exceeds a preset number threshold, the sequences in the genome of the virus to be analyzed will be clustered during alignment to obtain multiple sequence cluster sets. Each sequence cluster set includes at least one genome sequence. The sequence cluster set with the largest number of genome sequences is determined, and the longest genome sequence in the sequence cluster set with the largest number of genome sequences is determined as the representative genome sequence. The sequences in the genome of the virus to be analyzed are aligned to the representative genome sequence, and the single-site variation is calculated to generate a consistent sequence. Minimap2 alignment software can be used for this process.
[0076] For example, when N is 10000 bp and M is 100 sequences, if the genome sequence length of the virus to be analyzed is less than or equal to 10000 bp and the number of genome sequences is less than or equal to 100, sequence alignment is performed using MAFFT software to obtain a consistent sequence. If the genome sequence length of the virus to be analyzed is greater than 10000 bp and / or the number of genome sequences is greater than 100, when aligning the sequences in the genome of the virus to be analyzed, the sequences in the genome of the virus to be analyzed are clustered to obtain multiple sequence cluster sets, each of which includes at least one genome sequence; the sequence cluster set with the largest number of genome sequences is determined, and the longest genome sequence in the sequence cluster set with the largest number of genome sequences is determined as the representative genome sequence. The sequences in the genome of the virus to be analyzed are aligned to the representative genome sequence, and the single-site variation is calculated to generate a consistent sequence.
[0077] S230, based on the set sliding window, the consistency sequence of the viral genome to be analyzed is traversed, and a set of primer design index values are determined based on the sequence segment in each sliding window.
[0078] In this embodiment, the sliding window can be specifically understood as a window of preset length, which traverses the consistent sequence of the viral genome to be analyzed at a set step size to extract candidate sequence segments at different positions, and a set of primer design index values are determined based on the sequence segments in each sliding window.
[0079] Specifically, a sliding window is set on the homogeneous sequence of the viral genome to be analyzed and traversed, and a set of primer design index values are obtained based on the sequence segments within each sliding window. By setting a sliding window on the homogeneous sequence of the viral genome to be analyzed and traversing it, the comprehensiveness and uniformity of the primer design index values can be ensured.
[0080] Optionally, based on the above embodiments, the setting of the sliding window includes an F primer sliding window and an R primer sliding window, and correspondingly, the primer design index values include F primer design index values and R primer design index values.
[0081] In this embodiment, the F primer sliding window can be specifically understood as a dedicated sliding window adapted to the design requirements of F primers, used to extract candidate sequence segments that can serve as F primer binding sites, providing standardized input for deriving F primer design index values. Similarly, the R primer sliding window can be specifically understood as a sliding window adapted to the design requirements of R primers, used to extract candidate sequence segments that can serve as R primer binding sites, providing standardized input for deriving R primer design index values.
[0082] F-primer design metrics refer to the metrics established for F-primer design, including but not limited to length, GC content, Tm value, degeneracy bases, and structure. These metrics clearly define the acceptable design criteria for F-primers and ensure their synergistic adaptation with R-primers for PCR amplification requirements, serving as the core basis for accurate F-primer design. R-primer design metrics refer to the metrics established for R-primer design, including but not limited to length, GC content, Tm value, degeneracy bases, and structure. These metrics clearly define the acceptable design criteria for R-primers and ensure their synergistic adaptation with F-primers for PCR amplification requirements.
[0083] Optionally, based on the above embodiments, the method further includes: determining the sliding window range of the first F primer sliding window and the sliding window range of the first R primer sliding window based on a pre-set amplification product length range and primer length range; for non-first sliding windows, determining the sliding window range of non-first F primer sliding windows and non-first R primer sliding windows based on the amplification product length range, the primer length range, the overlap length of adjacent amplification products, and the end coordinates of the previous R primer sliding window.
[0084] In this embodiment, the pre-set amplification product length range can be specifically understood as the length range [L1, L2] of the PCR amplification target fragment preset according to requirements, where L1 is the minimum amplification length and L2 is the maximum amplification length. This constrains the design regions of the F primer sliding window and the R primer sliding window, ensuring that the fragment length amplified by the primer pair meets experimental detection standards, while also considering amplification efficiency and specificity. The primer length range can be specifically understood as the length range [P1, P2] of the F primer and R primer preset according to PCR amplification principles, primer binding specificity, and amplification efficiency requirements, where P1 is the minimum primer length and P2 is the maximum primer length. This limits the length of the candidate sequence segment truncated by the sliding window.
[0085] The sliding window range of the first F primer can be specifically understood as the range initially defined on the homogeneous sequence of the viral genome to be analyzed, based on the pre-set amplification product length range and primer length range, and adapted to the F primer design. This ensures that the extracted candidate sequence segment can form an amplification product of length within [L1, L2] with the candidate site of the first R primer sliding window. The sliding window range of the first R primer can also be understood as the range initially defined on the homogeneous sequence of the viral genome to be analyzed, based on the pre-set amplification product length range and primer length range, and adapted to the R primer design. This ensures that the extracted candidate sequence segment can form an amplification product of length within [L1, L2] with the candidate site of the first F primer sliding window.
[0086] The end coordinates of the previous R primer sliding window can be understood as the termination coordinates of the preceding R primer sliding window within the homogeneous sequence of the viral genome to be analyzed during continuous sliding window traversal. This is crucial for calculating the ranges of the non-first F primer sliding window and the R primer sliding window, ensuring that the amplification products of consecutive sliding windows are connected through a preset overlap length S, avoiding omissions in genome sequence coverage. The sliding window range of the non-first F primer sliding window, during continuous sliding window traversal, is a sliding window range on the homogeneous sequence of the viral genome to be analyzed, adapted to the F primer design, based on pre-set amplification product length range, primer length range, adjacent amplification product overlap length S, and the end coordinates r of the previous R primer sliding window. Similarly, the sliding window range of the non-first R primer sliding window can be understood as a sliding window range on the homogeneous sequence of the viral genome to be analyzed, adapted to the R primer design, based on pre-set amplification product length range, primer length range, adjacent amplification product overlap length S, and the end coordinates r of the previous R primer sliding window during continuous sliding window traversal.
[0087] Specifically, the amplified product length ranges from [L1, L2], the primer length ranges from [P1, P2], and the overlap length between adjacent amplified products is S. If it is the first sliding window, the sliding window range of the first F primer is [1, (L2 - L1) / 2 + P2], and the sliding window range of the first R primer is [L2 - ((L2 - L1) / 2 + P2), L2]. If it is not the first sliding window, and the end coordinate of the R primer in the previous sliding window is r, then the sliding window range of the current non-first F primer is [rS - ((L2 - L1) / 2 + P2), rS], and the sliding window range of the non-first R primer is [r - S + L2 - ((L2 - L1) / 2 + P2), r - S + L2]. By precisely calculating and adjusting the sliding window ranges of the F and R primers through quantitative parameters, continuous and complete coverage of the viral genome is achieved, ensuring the comprehensiveness and adaptability of primer design.
[0088] Optionally, based on the above embodiments, the method further includes: performing sequence alignment on the viral genome sequence to be analyzed to determine the mutation matrix and the consistency sequence; determining the mutation matrix and the consistency sequence; identifying high-variability regions based on the mutation matrix and / or the consistency sequence, wherein the mutation rate of the sites in the high-variability regions exceeds the mutation threshold and / or the consistency index is less than the consistency threshold; performing clustering processing on the sequence segments corresponding to the high-variability regions to obtain multiple cluster groups, wherein the mutation rate of the sites in each cluster group is less than the mutation threshold and the consistency index is greater than the consistency threshold; and determining a set of primer design index values for each cluster group.
[0089] In this embodiment, the variation matrix can be specifically understood as a matrix that presents site variation information after performing multiple sequence alignment on other genomic sequences of the virus genome to be analyzed. The matrix rows correspond to each non-representative genomic sequence participating in the alignment, and the columns correspond to each base site of the representative genomic sequence. Its core function is to systematically describe the variation distribution characteristics of all sequences in the virus genome to be analyzed at each site, providing a quantitative basis for the identification and clustering of highly variable regions.
[0090] Highly variable regions can be specifically understood as contiguous base regions in the viral genome sequence to be analyzed, which, after variation matrix analysis and consistency sequence feature determination, meet the criteria of having a site mutation rate exceeding a preset mutation threshold and / or a site consistency index less than a preset consistency threshold. The core characteristic is that the degree of variation in the base sequence within this region is significantly higher than in other regions of the genome. A single primer cannot cover all mutation types within this region; therefore, targeted primer design is required after clustering and grouping. The site mutation rate can be specifically understood as the proportion of all sequences within the viral genome sequence to be analyzed that exhibit base variation at a specific site. This is the core quantitative indicator for determining highly variable sites and thus identifying highly variable regions. The mutation threshold can be specifically understood as a pre-set value used to determine whether a single base site is a highly variable site. It is the quantitative boundary between site conservatism and highly variable sites and a key criterion for identifying highly variable regions. The consistency index can be understood as calculating the percentage of the most frequently occurring bases at a specific base site in the analyzed viral genome sequence across all genomic sequences. Its core function is to quantify the conservation of a single base site, corresponding to the site variation rate, and together they serve as the core quantitative basis for identifying highly variable sites and regions. The consistency threshold can be understood as a pre-set critical value used to determine the conservation of a single base site. Its core function is to define the quantitative boundary between site conservation and high variability, and together with the variation threshold, they serve as the key criterion for identifying highly variable regions. Clustering groups can be understood as multiple clusters obtained by clustering sequence segments corresponding to highly variable regions.
[0091] Specifically, when aligning sequences in the viral genome to be analyzed, if the sequence length is less than a preset length threshold or the number of genome sequences is less than a preset quantity threshold, MAFFT alignment software can be used. However, if the sequence length is greater than or equal to the preset length threshold or the number of genome sequences is greater than or equal to the preset quantity threshold, the MAFFT computation time will increase significantly. In such cases, a representative genome sequence of the viral genome to be analyzed can be selected, and other genome sequences of the viral genome to be analyzed can be aligned to the representative genome sequence using minimap2 alignment software. Based on the individual site variation, a variation matrix is constructed, and a consensus sequence is generated. Regions with consecutive sites where the variation rate exceeds the variation threshold and / or the consensus index is less than the consensus threshold are considered high-variability regions. The sequence segments corresponding to high-variability regions are clustered to obtain multiple cluster groups. Within each cluster group, the site variation rate is less than the variation threshold, and the consensus index is greater than the consensus threshold. For each cluster, a set of primer design index values is determined. This can be multiple F primer design index values corresponding to one R primer design index value, or multiple R primer design index values corresponding to one F primer design index value. Subsequently, primer pairs with multiple F primers corresponding to one R primer, or one F primer corresponding to multiple R primers, are generated.
[0092] S240, Based on the primer design index value, generate target viral primers adapted to the genome of the virus to be analyzed.
[0093] Optionally, based on the above embodiments, generating target viral primers adapted to the genome of the virus to be analyzed based on the primer design index values includes: generating multiple candidate primers based on the primer design index values; evaluating the candidate primers for primer dimer and hairpin structure; and determining the target viral primers based on the evaluation results.
[0094] In this embodiment, candidate primers can be specifically understood as primers that initially meet the requirements of the primer design index, selected from the candidate sequence range of the primer sliding window, but have not yet been verified for specificity, secondary structure, etc., and form the basis for screening target virus primers. Primer dimer evaluation can be specifically understood as a quantitative assessment of the base complementary pairing potential between each primer in the candidate primers. The core purpose is to determine whether the primer combination will form a stable primer dimer structure. The formation of primer dimers can lead to non-specific binding of primers in the PCR reaction, thereby reducing the amplification efficiency of the target fragment. Hairpin structure evaluation can be specifically understood as a quantitative assessment of the base complementary pairing potential of a single candidate primer. The core purpose is to determine whether the primer sequence will form a stable hairpin secondary structure due to local base complementarity. This structure can hinder the specific binding of the primer to the viral genome target sequence, leading to a decrease in PCR amplification efficiency or even amplification failure. The target virus primer can be specifically understood as the final primer sequence selected from the candidate primer set after performance verification such as primer dimer evaluation and hairpin structure evaluation. It is completely adapted to the characteristics of the genome sequence of the virus to be analyzed and meets the requirements of PCR amplification specificity and efficiency.
[0095] Specifically, multiple candidate primers are generated based on primer design index values. These candidate primers are then evaluated for primer dimer and hairpin structure. The target viral primer is determined based on the evaluation results. By evaluating the dimer and hairpin structure of the candidate primers, primers with structural defects can be accurately screened out, ensuring the amplification specificity and efficiency of the target viral primer.
[0096] Optionally, based on the above embodiments, generating target viral primers adapted to the viral genome to be analyzed based on the primer design index values includes: generating multiple candidate primers based on the primer design index values; determining a weighted penalty value for each candidate primer based on the difference between the primer design index value corresponding to each candidate primer and a preset optimal primer design index value; sorting the candidate primers based on the weighted penalty values of the candidate primers; and determining the target viral primers adapted to the viral genome to be analyzed.
[0097] In this embodiment, the preset optimal primer design index value can be specifically understood as a set of ideal standard values for each dimension of primer indicators, pre-set based on PCR amplification technical requirements and primer functional needs. This serves as a quantitative reference benchmark to measure the degree to which candidate primers conform to the expected optimal performance. The weighted penalty value of the candidate primer can be specifically understood as a quantitative score obtained by weighting the difference between the actual values of each dimension of the candidate primer indicators and the preset optimal primer design index value. Its core function is to comprehensively measure the degree to which candidate primers deviate from the optimal standard, providing an objective quantitative basis for primer ranking and screening.
[0098] Specifically, the basic indicators of candidate primers are calculated and the preset optimal primer design indicators are compared. The indicators may include, but are not limited to, the difference between the optimal primer length and the optimal GC content. A weighted penalty score is calculated and the primers are sorted from low to high according to the penalty score to determine the target viral primers that are suitable for the viral genome to be analyzed.
[0099] This embodiment of the disclosure obtains the genome of a virus to be analyzed; performs quality detection on the genome sequence in the genome to be analyzed, and removes genome sequences that do not meet the quality conditions; performs sequence alignment on the genome sequence to be analyzed to obtain the homogeneous sequence of the genome to be analyzed; traverses the homogeneous sequence of the genome to be analyzed based on a set sliding window, and determines a set of primer design index values based on the sequence segment in each sliding window, thereby realizing primer design for the entire viral genome, improving the automation of the primer design process, reducing errors caused by human intervention, and ensuring the full coverage of the viral genome by the target primer.
[0100] Based on the above embodiments, an optional example is provided, which can be used for scenarios involving the design of typing primers and whole-genome primers for target viruses.
[0101] The flowchart of the primer design pattern for virus typing is as follows: Figure 3As shown, genotyping primers for Coxsackievirus A10 (CA10), Coxsackievirus A16 (CA16), and Enterovirus A71 (EV71) were designed. First, the genome sequences of Coxsackievirus A10, Coxsackievirus A16, and Enterovirus 71 were downloaded from the NCBI Virus open-source database. To exclude fragmented and truncated low-quality sequences, the fragment length screening threshold for the viral genome sequences was set to 6500 bp. After filtering out redundant sequences and sequences with an N ratio of not less than 20%, high-quality sequences were obtained. The genomic sequences from multiple subtype genomic sequence sets are segmented to obtain multiple sequence segments, each of which generates a FastQ file. These sequence segments are then aligned using BWA with the representative genomic sequence of each viral subtype. The alignment results include the positional depth and base distribution matrix of the sequence segments. Conserved regions are screened based on an out-of-group depth threshold, resulting in continuous sequence segments that satisfy both the in-group and out-of-group depth thresholds. These continuous sequence segments are identified as conserved regions. The in-group conserved region is found to fall within gene VP1. Sequences are then aligned with these in-group conserved regions using BLASTN to extract the sequence set, followed by MAFFT alignment. Primers are then designed based on primer size, GC content, Tm value, etc., generating all candidate primers. Hairpin structures and primer dimer evaluations are performed to obtain the target viral primers. For example, a pair of primers was designed for EV71: the F primer is CTGCGAGTGCYTAYCARTGGTT; the R primer is ACYGAGAAMGTGCCCA.
[0102] The flowchart of primer design for the whole genome of the virus is as follows: Figure 4 As shown, the complete genome sequence of Coxsackievirus A10 was designed. The Coxsackievirus A10 sequence was downloaded from the open-source database NCBI Virus. The fragment length screening threshold for the viral genome sequence was set to 6500 bp, resulting in 487 sequences. After filtering out redundant sequences and sequences with an N ratio of at least 20%, 451 high-quality sequences remained. The amplification product fragment size was set to 1000-1200 bp, and a sliding window traversal was performed, resulting in the design of 24 primers in 8 regions. Furthermore, the viral genome contains hypervariable regions; designing only one primer would lead to excessive degenerate bases or insufficient coverage. The program automatically clustered the sequences based on a consistency score of 0.8, dividing them into two groups for separate design, ultimately resulting in two R primers, ensuring adequate coverage.
[0103] Figure 5This is a schematic diagram of a target virus primer generation device provided in an embodiment of this disclosure. This embodiment is applicable to the design of genotyping primers and multiplex primers targeting viral genomes. The device can be implemented using software and / or hardware, and can be integrated into any device that provides the function of generating target virus primers, such as… Figure 5 As shown, the device for obtaining the target virus primers specifically includes: a module 510 for obtaining the genome of the virus to be analyzed, a module 520 for obtaining primer design index values, and a module 530 for obtaining the target virus primers.
[0104] The viral genome acquisition module 510 is used to acquire the viral genome to be analyzed.
[0105] The primer design index value acquisition module 520 is used to analyze the genome of the virus to be analyzed based on the primer design pattern of the genome of the virus to be analyzed, and obtain the primer design index value.
[0106] The target virus primer acquisition module 530 is used to generate target virus primers adapted to the genome of the virus to be analyzed based on the primer design index values.
[0107] This embodiment of the disclosure obtains the genome of a virus to be analyzed; analyzes the genome based on the primer design pattern of the genome to be analyzed to obtain primer design index values; generates target viral primers adapted to the genome to be analyzed based on the primer design index values, and analyzes the viral genome through the primer design pattern to obtain primer design index values. This solves the problem of relying on manual methods for sequence quality control and design region selection, and realizes the automation, efficiency and accuracy of viral primer design.
[0108] Based on the above embodiments, optionally, the viral genome to be analyzed includes multiple viral subtype genomes, and the primer design mode includes a viral typing primer design mode, which is used to generate target viral primers corresponding to each viral subtype.
[0109] Based on the above embodiments, optionally, the primer design index value acquisition module 520 is used to: identify the representative genomic sequence of each viral subtype genome; segment the genomic sequences in the multiple viral subtype genome sequence set to obtain multiple sequence segments; compare the multiple sequence segments with the representative genomic sequence of each viral subtype to obtain the comparison result; determine the conserved region in each viral subtype genome based on the comparison result; and determine the primer design index value corresponding to each viral subtype genome based on the conserved region in each viral subtype genome.
[0110] Based on the above embodiments, optionally, the primer design index value acquisition module 520 is further configured to: perform clustering processing on each viral subtype genome sequence set to obtain clustering results, wherein the clustering results include at least one sequence cluster set, and each sequence cluster set includes at least one genome sequence; determine the sequence cluster set with the largest number of genome sequences, and determine the longest genome sequence in the sequence cluster set with the largest number of genome sequences as the representative genome sequence of the viral subtype.
[0111] Based on the above embodiments, optionally, the alignment result includes the positional depth of the sequence segment.
[0112] Based on the above embodiments, optionally, the primer design index value acquisition module 520 is further configured to: determine the intragroup depth of each sequence segment in the representative genome sequence of the viral subtype and the outgroup depth of each sequence segment in the representative genome sequence of other viral subtypes based on the alignment results; filter the intragroup depth of each sequence segment based on the intragroup depth threshold and the outgroup depth of each sequence segment based on the outgroup depth threshold to obtain continuous sequence segments that satisfy the intragroup depth threshold and the outgroup depth threshold, and determine the continuous sequence segments as the conserved region.
[0113] Optionally, based on the above embodiments, the alignment result may also include a base distribution matrix.
[0114] Based on the above embodiments, optionally, the primer design index value acquisition module 520 is further used to: perform intragroup consistency verification of the continuous sequence segment in the genome sequence set of its respective viral subtype based on the base distribution matrix.
[0115] Based on the above embodiments, optionally, the primer design index value acquisition module 520 is further configured to: extract the sequence segments of the genome sequence relative to the conserved region in each viral subtype genome sequence set to form a sequence set; perform multiple sequence alignment on multiple sequence segments in the sequence set to obtain multiple sequence alignment results; and determine the primer design index value corresponding to each viral subtype genome based on the multiple sequence alignment results.
[0116] Based on the above embodiments, optionally, the viral genome to be analyzed includes the whole viral genome, which includes multiple subtype genomes; the primer design mode includes a whole viral genome primer design mode, which is used to generate target viral primers corresponding to a set type of virus.
[0117] Based on the above embodiments, optionally, the primer design index value acquisition module 520 is further configured to: perform sequence alignment on the genome sequence of the virus to be analyzed to obtain the consistency sequence of the genome of the virus to be analyzed; traverse the consistency sequence of the genome of the virus to be analyzed based on a set sliding window, and determine a set of primer design index values based on the sequence segment in each sliding window.
[0118] Based on the above embodiments, optionally, the set sliding window includes an F primer sliding window and an R primer sliding window, and correspondingly, the primer design index values include F primer design index values and R primer design index values.
[0119] Based on the above embodiments, optionally, the primer design index value acquisition module 520 is further configured to: determine the sliding window range of the first F primer sliding window and the sliding window range of the first R primer sliding window based on the pre-set amplification product length range and primer length range; for non-first sliding windows, determine the sliding window range of non-first F primer sliding windows and non-first R primer sliding windows based on the amplification product length range, the primer length range, the overlap length of adjacent amplification products and the end coordinates of the previous R primer sliding window.
[0120] Based on the above embodiments, optionally, the primer design index value acquisition module 520 is further configured to: perform sequence alignment on the viral genome sequence to be analyzed to determine the mutation matrix and the consistency sequence; determine high-variability regions based on the mutation matrix and / or the consistency sequence, wherein the mutation rate of the sites in the high-variability regions exceeds the mutation threshold and / or the consistency index is less than the consistency threshold; perform clustering processing on the sequence segments corresponding to the high-variability regions to obtain multiple cluster groups, wherein the mutation rate of the sites in each cluster group is less than the mutation threshold and the consistency index is greater than the consistency threshold; and determine a set of primer design index values for each cluster group.
[0121] Based on the above embodiments, optionally, the target virus primer acquisition module 530 is used to: generate multiple candidate primers based on the primer design index values; evaluate the candidate primers for primer dimer and hairpin structure, and determine the target virus primer based on the evaluation results.
[0122] Based on the above embodiments, optionally, the target virus primer acquisition module 530 is further configured to: generate multiple candidate primers based on the primer design index value; determine the weighted penalty value of the candidate primer based on the difference between the primer design index value corresponding to each candidate primer and the preset optimal primer design index value; sort the candidate primers based on the weighted penalty value of the candidate primers; and determine the target virus primers adapted to the genome of the virus to be analyzed.
[0123] Optionally, based on the above embodiments, the device may further include a quality detection module, used to: perform quality detection on the genomic sequences in the viral genome to be analyzed, and remove genomic sequences that do not meet the quality conditions.
[0124] The above-described products can perform the methods provided in any embodiment of this disclosure, and have the corresponding functional modules and beneficial effects for performing the methods.
[0125] Figure 6 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0126] like Figure 6 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0127] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0128] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as the target virus primer generation method.
[0129] In some embodiments, the target virus primer generation method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the target virus primer generation method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the target virus primer generation method by any other suitable means (e.g., by means of firmware).
[0130] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0131] Computer programs used to implement the methods of this disclosure may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0132] In the context of this disclosure, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0133] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0134] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0135] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0136] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this disclosure can be achieved, and this is not limited herein.
[0137] This disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the target virus primer generation method according to any embodiment of this disclosure.
[0138] In implementing a computer program product, computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0139] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for generating primers for a target virus, characterized in that, include: Obtain the genome of the virus to be analyzed; Based on the primer design pattern of the viral genome to be analyzed, the viral genome to be analyzed is analyzed to obtain primer design index values; Based on the primer design index values, target viral primers adapted to the genome of the virus to be analyzed are generated.
2. The method according to claim 1, characterized in that, The viral genome to be analyzed includes multiple viral subtype genomes, and the primer design mode includes a viral typing primer design mode, which is used to generate target viral primers corresponding to each viral subtype. The genome of the virus to be analyzed was analyzed to obtain primer design index values, including: Identify the representative genome sequence for each viral subtype; The genome sequences in the multiple viral subtype genome sequence sets are segmented to obtain multiple sequence segments; The multiple sequence segments are compared with the representative genome sequence of each viral subtype to obtain the comparison results; Based on the alignment results, conserved regions in the genome of each viral subtype are determined, and primer design index values corresponding to each viral subtype genome are determined based on the conserved regions in the genome of each viral subtype.
3. The method according to claim 2, characterized in that, The representative genome sequence for each viral subtype is identified, including: Clustering is performed on each set of viral subtype genome sequences to obtain clustering results, wherein the clustering results include at least one sequence cluster set, and each sequence cluster set includes at least one genome sequence; A sequence cluster set with the largest number of genome sequences is determined, and the longest genome sequence in the sequence cluster set with the largest number of genome sequences is determined as the representative genome sequence of the viral subtype.
4. The method according to claim 2, characterized in that, The alignment results include the positional depth of the sequence segment; Based on the comparison results, conserved regions in the genome of each viral subtype were determined, including: Based on the alignment results, the intragroup depth of each sequence segment in the representative genome sequence of its respective viral subtype and the outgroup depth of each sequence segment in the representative genome sequences of other viral subtypes are determined. The intra-group depth of each sequence segment is filtered based on an intra-group depth threshold, and the extra-group depth of each sequence segment is filtered based on an extra-group depth threshold, to obtain continuous sequence segments that satisfy both the intra-group depth threshold and the extra-group depth threshold, and the continuous sequence segments are determined as the conservative region.
5. The method according to claim 4, characterized in that, The alignment results also include a base distribution matrix; The method further includes: Based on the base distribution matrix, intragroup consistency verification of the continuous sequence segment within the genome sequence set of its respective viral subtype is performed.
6. The method according to claim 2, characterized in that, Based on the conserved regions in the genome of each viral subtype, primer design index values corresponding to each viral subtype are determined, including: Extract the genomic sequence segments relative to the conserved regions from the genomic sequence set of each viral subtype to form a sequence set; Multiple sequence alignments are performed on multiple sequence segments in the sequence set to obtain the multiple sequence alignment results; Based on the results of the multiple sequence alignment, the primer design index value corresponding to each viral subtype is determined.
7. The method according to claim 1, characterized in that, The viral genome to be analyzed includes the whole viral genome; the primer design mode includes the whole viral genome primer design mode, which is used to generate target viral primers corresponding to a specified type of virus. The genome of the virus to be analyzed was analyzed to obtain primer design index values, including: Sequence alignment was performed on the genome sequence of the virus to be analyzed to obtain a consistent sequence of the genome of the virus to be analyzed; The consistency sequence of the viral genome to be analyzed is traversed by setting a sliding window, and a set of primer design index values are determined based on the sequence segment in each sliding window.
8. The method according to claim 7, characterized in that, The set sliding window includes an F primer sliding window and an R primer sliding window, and correspondingly, the primer design index values include F primer design index values and R primer design index values; The method further includes: The sliding window range of the first F primer and the sliding window range of the first R primer are determined based on the pre-set amplification product length range and primer length range. For non-first sliding windows, the sliding window range of non-first F primer sliding windows and non-first R primer sliding windows are determined based on the length range of the amplification product, the length range of the primer, the overlap length of adjacent amplification products, and the end coordinates of the previous R primer sliding window.
9. The method according to claim 7, characterized in that, The method further includes: Sequence alignment was performed on the genome sequence of the virus to be analyzed to determine the variation matrix and the sequence of consistency. High-variance regions are determined based on the mutation matrix and / or consistency sequence, wherein the mutation rate of points in the high-variance regions exceeds the mutation threshold and / or the consistency index is less than the consistency threshold. Clustering is performed on the sequence segments corresponding to the high-variability region to obtain multiple cluster groups. In each cluster group, the variation rate of the points is less than the variation threshold, and the consistency index is greater than the consistency threshold. For each cluster group, a set of primer design index values are determined.
10. The method according to any one of claims 1-9, characterized in that, Based on the primer design index values, target viral primers adapted to the genome of the virus to be analyzed are generated, including: Multiple candidate primers are generated based on the primer design index values; The candidate primers were evaluated for primer dimer and hairpin structure, and the target virus primers were determined based on the evaluation results.
11. The method according to any one of claims 1-9, characterized in that, Based on the primer design index values, target viral primers adapted to the genome of the virus to be analyzed are generated, including: Multiple candidate primers are generated based on the primer design index values; Based on the difference between the primer design index value corresponding to each candidate primer and the preset optimal primer design index value, the weighted penalty value of the candidate primer is determined. The candidate primers are then sorted based on the weighted penalty values to determine the target virus primers that are compatible with the viral genome to be analyzed.
12. The method according to any one of claims 1-9, characterized in that, The method further includes: The genomic sequences in the viral genome to be analyzed are subjected to quality testing, and genomic sequences that do not meet the quality requirements are removed.
13. A target virus primer generation device, characterized in that, include: The module for acquiring the genome of the virus to be analyzed is used to acquire the genome of the virus to be analyzed. The primer design index value acquisition module is used to analyze the genome of the virus to be analyzed based on the primer design pattern of the genome of the virus to be analyzed, and obtain the primer design index value. The target virus primer acquisition module is used to generate target virus primers adapted to the genome of the virus to be analyzed based on the primer design index values.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that are used to cause a processor to execute the method for generating target virus primers according to any one of claims 1-12.
15. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the target virus primer generation method according to any one of claims 1-12.