System and method for searching nucleic acid detection design area
By establishing a classifier model using CNN and analyzing the frequency of feature sequence fragments, the problem of finding high-retention primers and probes quickly in traditional methods is solved, achieving efficient and rapid primer design suitable for multiplex pathogen detection.
Patent Information
- Application Number
- CN202410777116.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-17
- Publication Date
- 2025-12-19
AI Technical Summary
Existing technologies struggle to quickly and effectively find highly preserved and specific primer-probe sets when designing nucleic acid detection for multiple pathogens. This is especially true when dealing with emerging infectious diseases or unpredictable pathogens. Traditional methods, such as multiple sequence alignment tools, are limited by sequence length and similarity, making it difficult to efficiently process genome sequences of multiple gene combinations.
A classifier model is built using a convolutional neural network (CNN). By analyzing the overlap characteristics and frequency of occurrence of feature sequence fragments, high-retention feature regions are identified. Primer probe sets are designed to replace traditional multiple sequence alignment methods and are suitable for primer region exploration of long gene sequences.
It enables rapid and efficient identification of primer design regions with high retention and specificity, with a computation speed more than 10 times faster than traditional methods. It is applicable to long genome sequences and does not require prior collection of gene information. It is suitable for novel pathogens or regions that circumvent patent protection.
Smart Images

Figure CN121171353A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to a system and method for finding a design region for nucleic acid detection, and in particular, to a system and method for finding a design region for nucleic acid detection. BACKGROUND
[0002] Due to the rapid development of globalization, there has been an increasing trend of emerging infectious diseases in recent years. The large-scale outbreak of coronavirus disease 2019 (COVID-19) has greatly affected the health and economic activities of the world in the past three years. Effective prevention and control of the spread of the epidemic relies on rapid and accurate detection methods. In addition to rapid screening, the nucleic acid of the target biomarker detected by PCR is the standard for diagnosis. In the face of unpredictable emerging infectious diseases, it is necessary to quickly identify the source of infection and accurately detect multiple disease targets for precision medicine. We urgently need an efficient design method for multiple target nucleic acid detection. SUMMARY
[0003] One embodiment of the present disclosure provides a system for finding a design region for nucleic acid detection, comprising a computer processor and a memory. The memory stores a plurality of computer program instructions which, when executed by the computer processor, cause the computer processor to implement steps comprising: inputting two groups of genomic sequences; training a classifier model using a machine learning algorithm, configured to distinguish between the two groups of genomic sequences, to obtain a plurality of characteristic sequence fragments; extending the plurality of characteristic sequence fragments into a plurality of characteristic regions by the characteristic sequence fragments having overlapping characteristics; and analyzing the frequency of occurrence of the characteristic sequence fragments in one of the characteristic regions to generate a region retention score, wherein the frequency of occurrence is positively correlated with the sequence retention of the one of the characteristic regions.
[0004] In some embodiments, the machine learning algorithm comprises a convolutional neural network (CNN), a long short-term memory (LSTM), a recurrent neural network (RNN), a generative adversarial network (GAN), a radial basis function network (RBFN), a multilayer perceptron (MLP), a self-organizing map (SOM), a deep belief network (DBN), a restricted Boltzmann machine (RBM), an auto-encoder, or a combination thereof.
[0005] In some embodiments, the classifier model is further configured to fragment and numerize the genomic sequences with a sliding window of a specific nucleotide length to obtain maximum convolution values.
[0006] In some embodiments, the classifier model is further configured to retrieve the feature sequence fragments of the genomic sequences with the maximum convolution values and the corresponding positions of the maximum convolution values.
[0007] In some embodiments, the specific nucleotide length comprises 17 to 23 base pairs (b.p. or mer).
[0008] In some embodiments, in the step of analyzing the frequency of occurrence of the feature sequence fragments occurring in one of the feature regions, the feature regions comprise a region retention score generated by summing the frequency of occurrence of the feature sequence fragments occurring in one of the feature regions and dividing by the length of the feature region.
[0009] In some embodiments, after generating the region retention scores, the steps further comprise generating a highest retention template sequence, wherein one of the genomic sequences in one of the groups of genomic sequences has the highest retention score when the sum of all region retention scores of the one genomic sequence is higher than the sum of all region retention scores of any other genomic sequence, and the genomic sequence with the highest score is the highest retention template sequence.
[0010] In some embodiments, after generating the highest retention template sequence, the steps further comprise designing primer pairs for the feature regions of the highest retention template sequence.
[0011] Another embodiment of the present disclosure provides a method for finding a nucleic acid detection design region, comprising: inputting two groups of genomic sequences; training a classifier model using a machine learning algorithm, configured to distinguish between the two groups of genomic sequences, to obtain a plurality of feature sequence segments; extending into a plurality of feature regions by the feature sequence segments having overlapping characteristics; and analyzing the frequency of occurrence of the feature sequence segments in which one of the feature regions occurs to generate a region retention score, wherein the frequency of occurrence is positively correlated with the sequence retention of the feature region. BRIEF DESCRIPTION OF DRAWINGS
[0012] The various aspects of the present disclosure will be most readily understood in view of the following detailed description when read in conjunction with the drawings. It should be noted that the various features may not be drawn to scale according to industry standard drafting practices. In fact, the dimensions of the various features can be arbitrarily increased or decreased for the sake of discussion clarity. In order that the disclosure herein may be more fully understood, the following description, taken in conjunction with the accompanying drawings, should be considered.
[0013] Figure 1 A flowchart showing primer design using CNN to find target regions of some embodiments of the present disclosure.
[0014] Figure 2 A line graph showing the feature region score (Score) and average retention (AvgCon) of ADV of some embodiments of the present disclosure.
[0015] Figure 3 A line graph showing the feature region score and average retention of RNV of some embodiments of the present disclosure.
[0016] REFERENCE NUMERALS
[0017] 100: method,
[0018] S110A, S110B, S120, S130, S140, S150: steps. DETAILED DESCRIPTION
[0019] In order to make the description of the present disclosure more detailed and complete, the following describes the illustrative description for the embodiments and specific examples of the present disclosure, but this is not the only form of implementation or use of the specific examples of the present disclosure. The embodiments disclosed below can be combined with each other or replaced by each other in a beneficial case, and other embodiments can be added in an embodiment without further description or illustration. In the following description, many specific details will be described in detail to enable the reader to fully understand the following embodiments. However, the embodiments of the present disclosure can also be practiced without such specific details.
[0020] In addition, spatially relative terms, such as "under", "below", "lower", "over", "upper" and the like, can be used herein for ease of description to describe one element or feature's relationship to another element(s) or feature(s) as illustrated in the figures. The spatially relative terms are intended to encompass different orientations of the device in use or operation in addition to the orientations depicted in the figures. The devices can be otherwise oriented (rotated 90 degrees or at other orientations) and the spatially relative descriptors used herein interpreted accordingly.
[0021] As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises", "comprising", "includes" and / or "including", when used herein, specify the presence of stated features, integers, steps, operations, elements, and / or components but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0022] Further, when a number or a numerical range is described as "about", "approximately", and the like, the terms are intended to encompass numbers that are within a reasonable range given the nature of the characteristics associated with the number. For example, a number or a numerical range encompasses a reasonable range that includes the described number, such as within + / - 10% of the described number, based on known manufacturing tolerances associated with manufacturing characteristics having characteristics associated with the number. Further, the present disclosure can refer to numbers and / or letters repeatedly in various instances. This repetition is for the purpose of simplicity and clarity and does not itself indicate a relationship between the various embodiments and / or configurations discussed.
[0023] The selection of appropriate primer / probe sets is one of the most important factors affecting the success of a PCR. A successful primer design is able to amplify a specific target nucleic acid sequence and the primer does not match other non-target sequences, while the probe needs to be designed within the region that is surrounded by the primer pair. For example, for the detection of COVID-19, the primer pair is able to detect SARS-CoV-2 sequences including its variants and avoid other closely related human coronaviruses. However, it is not an easy task to find the appropriate primer / probe sets, the primer needs to satisfy the sequence conservation of the target pathogen to cover various variants and at the same time satisfy the target specificity.
[0024] The conventional primer design process for multiplex pathogen starts with selecting target genes, which relies on prior knowledge of the homologous pathogens. After collecting the target genes of the same category of pathogens, the sequence similarity regions are aligned using multiple sequence alignment (MSA) tools, and then the highly conserved regions are manually selected. The primer design is then performed in the highly conserved regions using tools (e.g., Primer3 software), and the specificity of the candidate primers is then checked across the pathogens. The MSA and manual selection of conserved regions are the most time-consuming and limited by the sequence length. For example, a sequence set of 2000 sequences with a length of about 10,000 bp takes about 32 hours to calculate on a desktop workstation (6 cores, 32 GB memory). The above complex steps can also be completed using an automatic design process (e.g., Primer-BLAST), but due to the limitation of MSA, only a small group of similar pathogen target genes (e.g., a sequence set of 10 sequences with a length of about 15,000 bp) can be uploaded in one batch for primer searching, and then there is the labor-intensive work of integrating multiple batch results. When encountering novel pathogens and unable to obtain target genes from prior knowledge, or wanting to avoid other patent-protected target genes, it will be quite challenging to find primer design regions. Because MSA is limited by sequence length and similarity, it is usually not applicable to genome sequences composed of multiple genes.
[0025] In addition to the use of MSA, machine learning or deep learning can be applied to nucleic acid sequence analysis to obtain the advantage of alignment-free, such as classifying multiple closely related respiratory virus sequences, and the established convolutional neural network (CNN) classifier can accurately classify various variants of COVID. On the other hand, some other methods use deep learning to extract feature information of genome sequences. In some other ways, CNN is used to establish a classifier model for multiple human coronaviruses, and a specific length (22 bp) feature sequence is extracted as a primer candidate for COVID. In some other ways, two CNN classifiers are established to find forward primers and reverse primers, and a mixed strategy feature sequence extraction method can pick out primer sequences of variable length (18-25 bp).
[0026] The present disclosure uses CNN to establish a sequence classifier to find feature regions (100 bp or more) that meet the length of PCR amplicons, and recommend primer design regions with high conservation and specificity. Figure 1). The disclosed method 100 is more flexible, first find out the high conservation regions of the target pathogen, and then select and optimize the primer pairs in the region following the primer design rules. The disclosed method can replace MSA and does not require prior knowledge of the target gene, and can more efficiently perform multiple pathogen primer region exploration. In addition to viral sequences, the disclosed method can be applied to larger bacterial genome sequences (>1M bp), and we have built a web application system to implement the disclosed process.
[0027] The disclosed method 100 includes inputting two groups of genome sequences, such as pathogen 1 genome of step S110A and pathogen 2 genome of step S110B. Then, a machine learning algorithm is used to train a classifier model configured to distinguish between the two groups of genome sequences to obtain a plurality of characteristic sequence fragments; for example, step S120 CNN and feature extraction and step S130 to obtain characteristic sequences. Then, the occurrence frequency of the characteristic sequence fragments in which one of the characteristic regions appears is analyzed to generate a region conservation score, wherein the occurrence frequency is positively correlated with the sequence conservation of the characteristic region; for example, step S140 analyzes the compatibility and exclusivity of the characteristic regions. Finally, a primer sequence is designed as step S150.
[0028] The following examples and experimental examples are provided to more fully describe the disclosed system and method for finding a nucleic acid detection design region, but are not intended to limit the disclosure. The scope of protection of the disclosure is defined by the appended claims.
[0029] Embodiments
[0030] Although the following describes the disclosed method using a series of operations or steps, the order shown by these operations or steps should not be interpreted as limiting the disclosure. For example, certain operations or steps can be performed in a different order and / or simultaneously with other steps. In addition, not all shown operations, steps and / or features must be performed to implement embodiments of the disclosure. In addition, each operation or step described herein can include several sub-steps or actions.
[0031] The disclosed method uses a deep learning convolutional neural network (CNN) to build a model, find out the characteristic region sequence that matches the length of the PCR amplicon, and recommend a high-preservation and specific primer design region. One technical feature of the disclosure can be used to find a high-preservation region in a long genomic sequence, without the need to collect gene information in advance and select target genes. This advantage can more efficiently explore the primer target region of multiple pathogens, without developing novel pathogens for PCR or avoiding other patent protection areas. Commonly known art techniques often use multiple sequence alignment methods (MSA) to find preservation regions, but they are usually not suitable for similar sequence lengths and contents, and they are often not applicable to genomic sequences. Another advantage of the disclosed method is high efficiency, with a calculation speed that is more than 10 times faster than typical primer design processes (using Clustal Omega to perform the same environment comparison in the web application system we built).
[0032] The disclosed process can be roughly divided into two stages. In the first stage, we use a CNN to train a classifier model that can distinguish between multiple pathogen sequences. The model includes convolution layers, pooling layers, fully connected layers, and output layers, in which a 17mer length sliding window is used to fragment and numerically value the input nucleic acid sequence. After the classifier model is trained, the maximum convolution value and its corresponding position are used to retrieve the characteristic sequence fragments of the original input sequence. This stage can obtain a set of characteristic sequences representing each category of pathogens. We select a 17-length fragment because it meets the minimum primer length for PCR reagent development, and we then extend it to a characteristic region by the overlapping characteristics of the characteristic sequences. In the second stage, we analyze the frequency of characteristic sequences in the original training sequence set and infer the sequence preservation within the same pathogen category. After the characteristic sequences are screened for preservation and specificity, the overlapping characteristic sequences are integrated into a characteristic region and the region preservation score is calculated. Finally, high-preservation template sequences and characteristic region candidates are output.
[0033] Example 1
[0034] This example illustrates the process of finding multiple pathogen primer regions using CNN and characteristic region screening. We built a web application system to implement the entire process. The pathogen targets of this example are human adenovirus (ADV) and human rhinovirus (RNV). Assuming that the target gene of the unknown primer is unknown, we start from the viral whole genome sequence. Through CNN model construction and characteristic sequence extraction, the recommended primer design region of multiple pathogens is output. The detailed workflow is described in the following steps.
[0035] Step 1: Collect 737 ADV genome sequences (including B, C and E types) and 1275 RNV genome sequences (including A, B and C types) from NCBI nucleotide database (https: / / www.ncbi.nlm.nih.gov / nucleotide / ). The longest sequence is 36155 bp. We use the web interface to upload the FASTA files of the two classes of pathogen sequences (class0.fa and class1.fa, not shown) and send the job. After the job is done, the following steps 2 to 4 will be automatically run. The progress of each step in the pipeline can be viewed through the web interface, and the clustering status and conservation distribution of the input sequence database can also be observed. (not shown).
[0036] Step 2: The pipeline uses Tensorflow (v1.15) to build a CNN, and extracts 1 / 10 of the sequence set as a validation set in a hierarchical manner. The original sequence is slid by a window of 17mer length, and an input matrix space of about 13M bases is used. 12 convolution filters are set, and the maximum value is taken every 148 positions to become a pooling layer in order to reduce the dimension. There are a total of 196 nodes in the fully connected layer, and finally the sequence classification is output as ADV or RNV. The training process performs 100 epochs, and about 20 epochs can reach 100% prediction accuracy.
[0037] Step 3: Extract feature sequences from the 12 convolution filters of the trained model, a total of 168020 unique 17-mer short sequences.
[0038] Step 4: Perform efficient sequence alignment (using bowtie v1.2.3) on all short fragments, analyze the sequence alignment map (SAM), and calculate the frequency of each feature sequence in the set of viral genome sequences.
[0039] Step 5: The template sequence is selected to be the sequence with the highest conservation in the same group of pathogens. The application system calculates the average number of occurrences of the feature sequence as a reference index for template selection. In this example, we can select a Human adenovirus 7 sequence (type B, Accession: MW816049.1) as the ADV template sequence and a Rhinovirus 2 sequence (type A, Accession: MN749151.1) as the RNV template sequence according to the frequency ranking of the feature sequence and expert opinion.
[0040] Step 6: After the template sequences of each pathogen category are selected, the system calculates the score of the feature region (not shown) and recommends the appropriate primer design region accordingly. The feature region is integrated according to the overlapping characteristics of the feature fragments (17mers) and the score of the region is calculated (Score = total frequency of the feature fragments in the region / length of the region).
[0041] Step 7: The following describes the screening process of the feature region. Since the primer of the existing patent (US20040219517A1) is in the L3 gene of ADV, it can focus on the target gene and serve as a verification. We additionally perform traditional MSA and calculate the average conservation (the proportion of the most nucleic acids at this position after sequence alignment), and it can be found that the feature region score (Score) and the average conservation (AvgCon) are consistent at high values. In other words, compared with the traditional MSA method, the disclosed method is consistent with the ADV sequence positions 17500, 18300, and 20600 (Table 1), among which the feature region of position 20600 is the longest and consistent with the primer region of the aforementioned patent. The feature region shows high conservation in the traditional MSA of the ADV sequence set, and only a few positions have inter-species differences (not shown). Figure 2 , Table 1) has a consistent feature region score (Score) and average conservation (AvgCon) at high values.
[0042] Table 1
[0043]
[0044]
[0045] Step 8: The feature region score (Score) and average conservation (AvgCon) of RNV are both high at the first 650 positions of the sequence ( Figure 3 ). Other positions of the RNV sequence are not suitable for primer design due to high inter-species divergence, and this result also conforms to the RNV primer design region (only present in the 5'UTR region) in the literature.
[0046] Step 9: Regarding the comparison of computational efficiency, the present embodiment takes about 285 minutes to complete on a desktop workstation (6 computing cores, 32 GB of memory); compared with the MSA tool (Clustal Omega v1.2.1), which takes about 1937 minutes to execute on the full-length ADV and RNV sequence set (a total of 1942 pens), and due to the uneven similarity of the sequences, additional sequence alignment adjustment is required for subsequent multiple sequence alignment.
[0047] Example 2
[0048] This example illustrates the process of finding multiple pathogen primer regions using CNN and feature region screening. We built a web application system to perform the entire process. The pathogen target of this example is Gemella morbillorum (GM). Currently, there is no PCR-related literature found for GM bacteria. We started from its genome sequence to find suitable primer regions. Through CNN model construction and feature sequence extraction, we produced recommendations for multiple pathogen primer design regions. The detailed workflow is as follows.
[0049] Step 1: We collected 3 GM bacteria genome sequences and 15 Gemella genus closely related species (non-GM bacteria) genome sequences from the NCBI nucleotide sequence database (https: / / www.ncbi.nlm.nih.gov / nucleotide / ). The longest sequence is more than 1.9 million bp. We used the web interface to upload the FASTA file of the two categories of bacterial sequences and sent the work.
[0050] Step 2: The CNN modeling step and feature region screening process are as described in Example 1.
[0051] Step 3: We found feature sequences that can distinguish the target GM bacteria from its highly related species using the above model. We screened 5 longer (length greater than 200 bp) feature regions from the 1.7 million bp GM bacteria whole genome sequence template (Accession: CP046314) (Table 2). For example, the 207 bp long feature region Reg_ID:325 is completely consistent among the 3 GM species strains. Further Blast search of the NCBI nr sequence database using this feature region sequence showed high specificity (Table 3).
[0052] Table 2
[0053] Feature region Start End Length Number of primer pairs candidates (by Primer3) 325 19611 19818 207 3 1944 118600 1188645 201 0 3116 188420 188645 225 1 3845 230499 230714 215 0 5568 1013436 1013636 200 5
[0054] Table 3
[0055]
[0056] Step 4: Within these 5 feature regions, we used Primer3 software to design primers and found 9 sets of candidate primer and probe pairs (Table 2), and checked them through the default parameters of Primer3 software.
[0057] Step 5: This example shows that this process can be applied to finding primer regions for novel pathogenic bacteria (single bacterial genome length can be more than 1 million, and unknown target gene), without relying on multiple sequence alignment, we can efficiently find suitable primer design regions.
[0058] In some embodiments of the present disclosure, the technical features of the CNN find multiple pathogen feature regions, which have the following advantages: 1. The sequence alignment step is eliminated, and long genomic sequence sets can be searched for primer design regions that meet the requirements of both retention and specificity; 2. No pre-selection of target genes is required, and there is a special advantage for novel primer region recommendations.
[0059] In some embodiments of the present disclosure, the technical features of the high-retention feature region screening process have the following advantages: 1. The feature region is integrated from the overlapping characteristics of feature fragments (17mers), and the feature region score is calculated by the following formula: Score = total frequency of feature fragments in the region / region length; 2. The feature region sorted by score is consistent with the high-retention region of MSA (see step 7 of Example 1), indicating that the technical features can replace MSA to find sequence retention regions; 3. It can handle long sequences and low-similarity sequences that cannot be aligned by MSA.
[0060] In some embodiments of the present disclosure, the high-retention feature region is found faster: 1. The calculation speed of the present method is more than 10 times faster than the typical primer design process based on MSA (comparison of the same sequence set in the same execution environment, see step 9 of Example 1); 2. It can be applied to long sequences that cannot be run by MSA, such as bacterial sequence sets with a total length of more than 50 million.
[0061] Although the present disclosure has been disclosed in the embodiments as above, it is not intended to limit the present disclosure, and any person skilled in the art can make various changes and modifications without departing from the spirit and scope of the present disclosure, and therefore the protection scope of the present disclosure shall be subject to the appended patent claims.
Claims
1. A system for finding a design region for nucleic acid detection, wherein, A computer processor and a memory storing computer program instructions that, when executed by the computer processor, cause the computer processor to implement steps comprising: inputting two groups of genomic sequences; training a classifier model configured to distinguish between the two groups of genomic sequences using a machine learning algorithm to obtain feature sequence segments; extending the feature sequence segments into feature regions by virtue of the feature sequence segments having overlapping characteristics; and analyzing a frequency of occurrence of the feature sequence segments occurring in one of the feature regions to generate a region retention score, wherein the frequency of occurrence is positively correlated with a sequence retention of the one of the feature regions.
2. The system of claim 1, wherein, The machine learning algorithm comprises a convolutional neural network, a long short-term memory network, a recurrent neural network, a generative adversarial network, a radial basis function network, a multilayer perceptron, a self-organizing map, a deep belief network, a restricted Boltzmann machine, an autoencoder, or a combination thereof.
3. The system of claim 1, wherein, The classifier model is further configured to fragment and numerate the genomic sequences using a sliding window of a specific nucleotide length to obtain a maximum convolution value.
4. The system of claim 3, wherein, The classifier model is further configured to retrieve the feature sequence segments of the genomic sequences using the maximum convolution value and a corresponding position of the maximum convolution value.
5. The system of claim 3, wherein, The specific nucleotide length comprises 17 to 23 bases.
6. The system of claim 1, wherein, In the step of analyzing the frequency of occurrence of the feature sequence segments occurring in one of the feature regions, the region retention score is generated by summing the frequency of occurrence of the feature sequence segments occurring in the one of the feature regions and dividing the sum by a region length of the one of the feature regions.
7. The system of claim 1, wherein, After generating the region retention score, the steps further comprise: generating a highest retention template sequence, wherein one of the genomic sequences in one of the two groups of genomic sequences has a highest score as the highest retention template sequence when a sum of all region retention scores of the one of the genomic sequences is higher than a sum of all region retention scores of any other genomic sequence.
8. The system of claim 7, wherein, After generating the highest retention template sequence, the steps further comprise: designing primer pairs for the feature regions of the highest retention template sequence.
9. A method of finding a design region for nucleic acid detection, wherein, The steps comprise: inputting two groups of genomic sequences; training a classifier model configured to distinguish between the two groups of genomic sequences using a machine learning algorithm to obtain feature sequence segments; extending the feature sequence segments into feature regions by virtue of the feature sequence segments having overlapping characteristics; and analyzing a frequency of occurrence of the feature sequence segments occurring in one of the feature regions to generate a region retention score, wherein the frequency of occurrence is positively correlated with a sequence retention of the one of the feature regions. The machine learning algorithm comprises a convolutional neural network, a long short-term memory network, a recurrent neural network, a generative adversarial network, a radial basis function network, a multilayer perceptron, a self-organizing map, a deep belief network, a restricted Boltzmann machine, an autoencoder, or a combination thereof.
10. The method of claim 9, wherein, 11. The method of claim 9, wherein, The classifier model is further configured to fragment and numerate the plurality of genomic sequences with a sliding window of a specific nucleotide length to obtain a maximum convolution value.
12. The method of claim 11, wherein, The classifier model is further configured to retrieve the feature sequence fragments of the plurality of genomic sequences using the maximum convolution value and a corresponding position of the maximum convolution value.
13. The method of claim 11, wherein, The specific nucleotide length comprises 17 to 23 bases.
14. The method of claim 9, wherein, In the step of analyzing the occurrence frequency of the feature sequence fragments occurring in one of the feature regions, the feature regions are analyzed by summing the occurrence frequency of the feature sequence fragments occurring in one of the feature regions and dividing the summed occurrence frequency by a region length of the one of the feature regions to generate a region retention score.
15. The method of claim 9, wherein, After the step of generating the region retention score, the method further comprises: generating a highest retention template sequence, wherein one of the genomic sequences in one of the two groups of genomic sequences has a highest score as the highest retention template sequence when a sum of all region retention scores of the one of the genomic sequences is higher than a sum of all region retention scores of any other genomic sequence.
16. The method of claim 15, wherein, After the step of generating the highest retention template sequence, the method further comprises: designing primer pairs for the feature regions of the highest retention template sequence.
Citation Information
Patent Citations
Methods for rapid identification of pathogens in humans and animals
US20040219517A1