Computer-implemented method for providing a nucleic acid sequence data set for the design of oligonucleotides
Patent Information
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-08-31
- Publication Date
- 2026-08-12
Smart Images

Figure 112023016115340-PCT00003_ABST
Abstract
Description
Technology Field
[0001] Cross-reference regarding related applications
[0002] This patent application claims priority to Korean Patent Application No. 2020-0110636 filed with the Korean Intellectual Property Office on August 31, 2020, and the disclosures of said patent applications are incorporated herein by reference.
[0003] Technology field
[0004] The present invention relates to a computer-implemented method for providing a nucleic acid sequence data set for designing oligonucleotides used to detect target nucleic acid molecules of an organism of interest. Background Technology
[0005] The 21st-century healthcare paradigm is transforming from the era of public health and disease treatment to an era of extending healthy life expectancy through disease prevention and management. In line with this global trend shifting from a focus on therapeutic medicine to preventive medicine, in vitro diagnostics ( In Vitro The demand for diagnostics (IVD) is increasing significantly. Global population aging and the emergence of novel viruses are contributing factors to the growth of the in vitro diagnostics market. Furthermore, as treatment methods shift toward personalized medicine, the scope of conducting in vitro diagnostics prior to determining treatment or prescriptions is expanding.
[0006] Molecular diagnostics is the fastest-growing sector in the in vitro diagnostics (IVD) industry and is key to patient care continuity. Compared to other platforms with overlapping disease portfolios, molecular diagnostics offers strengths such as superior test precision, miniaturization capabilities, and rapid processing times. Driven by these strengths of molecular diagnostic technology, traditional chemical and immunological general diagnostic tests are gradually being replaced by molecular diagnostic tests.
[0007] Representative technologies used in the field of molecular diagnostics include Polymerase Chain Reaction (PCR), Next-Generation Sequencing (NGS), microarray, or fluorescent molecular conjugation. in situ There are hybridization, etc.
[0008] A nucleic acid amplification method known as polymerase chain reaction (hereinafter referred to as “PCR”) comprises a repeated cycle of denaturation of double-stranded DNA, annealing of oligonucleotide primers into a DNA template, and primer extension by DNA polymerase (Mullis et al., U.S. Patents No. 4,683,195, 4,683,202 and 4,800,159; Saiki et al., (1985) Science 230, 1350-1354).
[0009] PCR-based technologies are widely used not only for the amplification of target DNA sequences but also for scientific applications and methods in the fields of biology and medical research. Examples include the detection of target sequences, reverse transcriptase PCR (RT-PCR), differential display PCR (DD-PCR), PCR-based cloning of known or unknown genes, rapid amplification of cDNA ends (RACE), random priming PCR (AP-PCR), multiplex PCR, SNP genome typing, and PCR-based genome analysis.
[0010] These molecular diagnostic technologies analyze pathogens or risk factors for disease by confirming the presence or sequence of target nucleic acid molecules within a sample, and in most cases, the sequence of the target nucleic acid molecule is selectively amplified for analysis. To perform such molecular diagnostics, the design of the oligonucleotides used for the amplification and detection of target nucleic acid molecules is important.
[0011] For the detection of target nucleic acid molecules, the oligonucleotides (probes and / or primers) used must possess appropriate specificity and detectability, be suitable for the specific detection method, and meet the conditions set by the analyst. Therefore, the design of oligonucleotides tailored to the analytical purpose is crucial.
[0012] Whether the genes of higher animals such as humans and mammals, or those of various bacteria and viruses classified as pathogens, most target nucleic acid sequences contain sequence variations among individuals. In particular, RNA viruses are well known to possess high sequence variability (genetic diversity). To detect target nucleic acid molecules with genetic diversity with appropriate coverage, a more sophisticated design of oligonucleotides is required.
[0013] Various attempts have been made to design oligonucleotides for detecting target nucleic acid molecules with genetic diversity. A conventional method for designing such oligonucleotides is to identify conserved regions in multiple target nucleic acid molecules with genetic diversity and design oligonucleotides that hybridize to these regions (Wang, D et al., Proc. Natl Acad. Sci. USA, 99:15687-15692(2002)).
[0014] Designing oligonucleotides for conserved sites required the collection of target nucleic acid sequences containing numerous sequence variations classified as the same type. While existing methods initiated new attempts regarding the processing of various target nucleic acid sequence data for target nucleic acid molecules and the subsequent discovery of conserved sites, there had been no methodological progress in the collection and processing of homologous sequences provided for oligonucleotide design; instead, the process still relied on the personal knowledge and experience of the researchers collecting the sequences.
[0015] When using this manual method of sequence collection, the specificity of the designed oligonucleotide for the target nucleic acid sequence and the coverage of the target nucleic acid sequence detectable by the oligonucleotide are limited depending on the researcher's capabilities, and the development time for sequence collection increases.
[0016] To address these problems, it was necessary to develop a new automated method for effectively collecting target nucleic acid sequence data for target nucleic acid molecules.
[0017] Meanwhile, FIG. 1 is a flowchart of a process for providing a target nucleic acid sequence data set for a target nucleic acid molecule according to a method previously filed by the applicant (International Publication No. WO2019 / 212238). According to the flowchart of FIG. 1, nucleic acid sequence data is collected using keywords such as the name of the target nucleic acid molecule, the collected nucleic acid sequence data is aligned according to the length of the sequence, the longest sequence is determined as the representative sequence, nucleic acid sequence data having homology greater than a predetermined value with the representative sequence are grouped, and then the nucleic acid sequence data collected with the keywords and the nucleic acid sequence data having homology with the representative sequence are provided as a target nucleic acid sequence data set for the target nucleic acid molecule. The result of aligning a plurality of target nucleic acid sequences of the target nucleic acid sequence data set is shown in FIG. 2.
[0018] As can be seen in Figure 2 above, when a nucleic acid sequence data set for a target nucleic acid molecule was provided using the conventional method, the number of representative sequences was 25, and it was confirmed that alignment was not properly performed due to differences in homology between the sequences collected from the representative sequences. As a result, there was a problem in that unnecessary time was spent by analysts reviewing the aligned nucleic acid sequences.
[0019] Accordingly, the inventors recognized the need to develop a method for designing an oligonucleotide in which a plurality of target nucleic acid sequences for a target nucleic acid molecule are collected without omission and an alignment result can be properly formed so that the collected plurality of target nucleic acid sequences can be used for the design of the oligonucleotide, in order to design an oligonucleotide used to detect a target nucleic acid molecule of an organism of interest.
[0020] Throughout this specification, numerous cited literature and patent literature are referenced and their citations are indicated. The disclosures of the cited literature and patents are incorporated by reference into this specification in their entirety to more clearly explain the state of the art to which the present invention pertains and the content of the present invention. The problem to be solved
[0021] The inventors have endeavored to develop a computer-implemented method capable of effectively providing a nucleic acid sequence data set for use in designing oligonucleotides used for amplification or detection of target nucleic acid molecules, while overcoming the problems of the aforementioned conventional method. As a result, the inventors have completed the present invention by aligning nucleic acid sequence data collected from synonyms of target nucleic acid molecules according to taxonomic names and / or taxonomic IDs, selecting taxonomic representative sequences from among nucleic acid sequence data having the same taxonomic name and / or taxonomic ID, grouping the selected taxonomic representative sequences according to homology, selecting a group representative sequence from each group, and providing nucleic acid sequence data having homology of a predetermined value or greater with the group representative sequence, thereby confirming that the alignment results of multiple nucleic acid sequences are well formed to the extent that oligonucleotides can be designed.
[0022] Accordingly, the object of the present invention is to provide a computer-implemented method for providing a nucleic acid sequence data set for designing oligonucleotides used to detect target nucleic acid molecules of an organism of interest.
[0023] Another object of the present invention is to provide a computer-readable recording medium comprising instructions for implementing a processor for executing a method for providing a nucleic acid sequence data set for designing oligonucleotides used to detect target nucleic acid molecules of an organism of interest.
[0024] Other objects and advantages of the present invention will become more apparent from the following embodiments, claims, and drawings. means of solving the problem
[0025] According to one aspect of the present invention, the present invention provides a computer-implemented method for providing a nucleic acid sequence data set for designing oligonucleotides used to detect target nucleic acid molecules of an organism of interest, comprising the following steps:
[0026] (a) receiving the name of a target nucleic acid molecule and the name of an organism of interest, and retrieving synonyms of the organism of interest for the target nucleic acid molecule;
[0027] (b) a step of collecting nucleic acid sequence data included in nucleic acid records; each of the nucleic acid records relates to the organism of interest and includes a descriptor in which the name of the target nucleic acid molecule and at least one of the collected synonyms are listed,
[0028] (c) sorting the collected nucleic acid sequence data according to taxonomic name and / or taxonomic ID and selecting a taxonomic representative sequence from among the nucleic acid sequence data having the same taxonomic name and / or taxonomic ID;
[0029] (d) a step of grouping the selected taxonomic representative sequences according to homology and selecting a group representative sequence from each group; and
[0030] (e) A step of collecting nucleic acid sequence data having homology greater than a predetermined value with the representative sequence of the group above and providing it as a nucleic acid sequence data set for designing the oligonucleotide.
[0031] The inventors have endeavored to develop a computer-implemented method capable of overcoming the problems of the aforementioned conventional method and effectively providing a nucleic acid sequence data set for use in designing oligonucleotides used for amplification or detection of target nucleic acid molecules. As a result, the inventors aligned nucleic acid sequence data collected from synonyms of target nucleic acid molecules according to taxonomic names and / or taxonomic IDs, selected taxonomic representative sequences from among the nucleic acid sequence data having the same taxonomic name and / or taxonomic ID, grouped the selected taxonomic representative sequences according to homology, selected a group representative sequence from each group, and provided nucleic acid sequence data having homology of a predetermined value or greater with the group representative sequence, thereby confirming that the alignment results of multiple nucleic acid sequences were well formed to the extent that oligonucleotides could be designed.
[0032] In this specification, the term “organism of interest” refers to an organism containing a target nucleic acid molecule to be amplified or detected using an oligonucleotide (e.g., a primer or probe).
[0033] In this specification, the term “organism” means an organism belonging to a biological classification system, e.g., kingdom, phylum, class, order, family, genus, species, subspecies, variety, variant, subtype, genotype, sirotype, strain, isolate, or cultivar. An organism is, for example, a prokaryotic cell (e.g., Mycoplasma pneumoniae, Chlamydophila pneumoniae, Legionella pneumophila, Haemophilus influenzae, Streptococcus pneumoniae, Bordetella pertussis, Bordetella parapertussis, Neisseria meningitidis, Listeria monocytogenes, Streptococcus agalactiae, Campylobacter, Clostridium difficile, Clostridium perfringens, Salmonella, Escherichia coli, Shigella, Vibrio, Yersinia enterocolitica, Aeromonas, Chlamydia trachomatis, Neisseria gonorrhoeae, Trichomonas vaginalis, Mycoplasma hominis, Mycoplasma genitalium, Ureaplasma urealyticum, Ureaplasma parvum, Mycobacterium tuberculosis It includes ), eukaryotic cells (e.g., protozoa and parasites, fungi, yeast, higher plants, lower animals, and higher animals including mammals and humans), viruses, or viroids. Examples of parasites among the eukaryotic cells are Giardia lamblia, Entamoeba histolytica, Cryptosporidium, Blastocystis hominis, Dientamoeba fragilis, Cyclospora cayetanensisIncludes. Examples of the above viruses include influenza A virus (Flu A), influenza B virus (Flu B), respiratory syncytial virus A (RSV A), respiratory syncytial virus B (RSV B), parainfluenza virus 1 (PIV 1), parainfluenza virus 2 (PIV 2), parainfluenza virus 3 (PIV 3), parainfluenza virus 4 (PIV 4), metapneumovirus (MPV), human enterovirus (HEV), human bocavirus (HBoV), human rhinovirus (HRV), coronavirus, and adenovirus; and norovirus, rotavirus, adenovirus, astrovirus, and sapovirus that cause gastrointestinal diseases. In addition, examples of the above viruses include HPV (human papillomavirus), MERS-CoV (Middle East respiratory syndrome-related coronavirus), dengue virus, HSV (Herpes simplex virus), HHV (Human herpes virus), EMV (Epstein-Barr virus), VZV (Varicella zoster virus), CMV (Cytomegalovirus), HIV, hepatitis virus, and polio virus.
[0034] In this specification, the terms “target nucleic acid molecule,” “target molecule,” or “target nucleic acid” refer to nucleotide molecules within an organism to be detected. Target nucleic acid molecules are generally given specific names and include the entire genome and all nucleotide molecules constituting the genome (e.g., genes, pseudogenes, non-coding sequence molecules, untranslated regions, and parts of the genome). Target nucleic acid molecules include, for example, nucleic acids of the organism.
[0035] In this specification, a target nucleic acid molecule may refer to the entire nucleic acid molecule to be detected or a partial region of the nucleic acid molecule. In this specification, a target nucleic acid molecule may refer to a functional unit among nucleic acid molecules. The functional unit may be a gene. A gene refers to a physical and functional unit of genetic information composed of DNA or RNA. The gene includes both regions that encode proteins and regions that do not encode proteins. In this specification, the term “target gene” is a term used interchangeably with the target nucleic acid molecule when the target nucleic acid molecule refers to a gene portion that is a functional unit among physical nucleic acid molecules.
[0036] In this specification, the term “detection” means a measurement that provides a qualitative or quantitative indication of the presence or absence of a target nucleic acid molecule. Such detection includes identification, determination, or analysis.
[0037] As used herein, the term “oligonucleotide” refers to a linear oligomer of natural or modified monomers or linkages, comprising deoxyribonucleotides and ribonucleotides, capable of specifically hybridizing to a target nucleic acid sequence, and is naturally occurring or artificially synthesized. Oligonucleotides are particularly single-stranded for maximum efficiency in hybridization. Specifically, oligonucleotides are oligodeoxyribonucleotides. Oligonucleotides used in the present invention may include naturally occurring dNMPs (i.e., dAMP, dGMP, dCMP, and dTMP), nucleotide analogs, or derivatives. Additionally, oligonucleotides may also include ribonucleotides. For example, in the present invention, oligonucleotides are backbone-modified nucleotides, such as peptide nucleic acid (PNA) (M. Egholm et al., Nature, 365:566-568 (1993)), Locked Nucleic Acid (LNA) (WO1999 / 014226), Bridged Nucleic Acid (BNA) (WO2005 / 021570), phosphorothioate DNA, phosphorodithioate DNA, phosphoroamidate DNA, amide-linked DNA, MMI-linked DNA, 2'-O-methyl RNA, alpha-DNA and methylphosphonate DNA, sugar-modified nucleotides e.g., 2'-O-methyl RNA, 2'-fluoro RNA, 2'-amino RNA, 2'-O-alkyl DNA, 2'-O-allyl DNA, 2'-O-alkynyl DNA, hexose DNA, pyranosyl RNA and anhydrohexitol DNA, and nucleotides having base modifications e.g., C-5 substituted Pyrimidines (substituents include fluoro-, bromo-, chloro-, iodo-, methyl-, ethyl-, vinyl-, formyl-, ethityl-, propynyl-, alkynyl-, thiazoryl-, imidazoryl-, pyridyl-), 7-deazpurines having C-7 substituents (substituents include fluoro-, bromo-, chloro-, iodo-, methyl-, ethyl-, vinyl-, formyl-, alkynyl-, alkenyl-, thiazoryl-, imidazoryl-, pyridyl-), inosines, and diaminopurines may be included. In particular, as used herein, the term “oligonucleotide” refers to a single strand composed of deoxyribonucleotides. The term “oligonucleotide” includes oligonucleotides that hybridize with cleavage fragments occurring dependently on the target nucleic acid sequence.
[0038] According to one embodiment of the present invention, the oligonucleotide is a primer and / or probe.
[0039] As used herein, the term “primer” refers to an oligonucleotide that can act as an initiator for synthesis under conditions in which the synthesis of a primer extension product complementary to a nucleic acid strand (template) is induced, namely, the presence of a polymerase such as a nucleotide and DNA polymerase, and conditions of suitable temperature and pH. The primer must be sufficiently long to prime the synthesis of the extension product in the presence of the polymerase. The suitable length of the primer is determined by a plurality of factors, including, for example, temperature, application, and the source of the primer.
[0040] As used herein, the term “probe” refers to a single-stranded nucleic acid molecule comprising a region or regions complementary to a target nucleic acid sequence. Additionally, the probe may include a label capable of generating a signal for target detection.
[0041] The above oligonucleotides may have conventional primer and probe structures composed of sequences that hybridize to a target nucleic acid sequence. Alternatively, the structure of the above oligonucleotides may be modified to have a unique structure. For example, the above oligonucleotides may have the structure of a scorpion primer, a molecular beacon probe, a sunrise primer, a high beacon probe, a tagging probe, a DPO primer or probe (WO 2006 / 095981), and a PTO probe (Ref. WO 2012 / 096523).
[0042] The above oligonucleotides may be modified oligonucleotides, such as degenerate base-containing oligonucleotides and / or universal base-containing oligonucleotides, in which degenerate bases and / or universal bases are introduced into conventional primers or probes. As used herein, the terms “conventional primer,” “conventional probe,” and “conventional oligonucleotide” refer to conventional primers, probes, and oligonucleotides in which degenerate bases or non-natural bases are not introduced. According to one embodiment of the present invention, the degenerate base-containing oligonucleotide or universal base-containing oligonucleotide is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% of the unmodified oligonucleotide. According to one embodiment of the present invention, the range of the number of degenerate bases or universal bases introduced into the conventional oligonucleotide is specifically 7 or fewer, 5 or fewer, 4 or fewer, 3 or fewer, or 2 or fewer. Alternatively, the usage ratio of degenerate bases and / or universal bases introduced into the conventional oligonucleotide is specifically 25% or less, 20% or less, 18% or less, 16% or less, 14% or less, 12% or less, 10% or less, 8% or less, or 6% or less. The usage ratio of degenerate bases or universal bases represents the ratio of degenerate bases or universal bases to the total nucleotides of the oligonucleotide into which the degenerate bases or universal bases have been introduced. The degenerate bases include the following various degenerate bases known in the art: R: A or G; Y: C or T; S: G or C; W: A or T; K: G or T; M: A or C; B: C, G, or T; D: A, G or T; H: A, C or T; V: A, C or G; N: A, C, G or T.The above universal bases include the following various universal bases known in the art: deoxyinosine, inosine, 7-diaza-2'-deoxyinosine, 2-aza-2'-deoxyinosine, 2'-OMe inosine, 2'-F inosine, deoxy 3-nitropyrrole, 3-nitropyrrole, 2'-OMe 3-nitropyrrole, 2'-F 3-nitropyrrole, 1-(2'-deoxy-beta-D-ribofuranosyl)-3-nitropyrrole, deoxy 5-nitropyrrole, 5-nitroindole, 2'-OMe 5-nitroindole, 2'-F 5-nitroindole, deoxy 4-nitrobenzimidazole, 4-nitrobenzimidazole, deoxy 4-aminobenzimidazole, 4-aminobenzimidazole, deoxy Nebularin, 2'-F Nebularin, 2'-F 4-Nitrobenzimidazole, PNA-5-Introindole, PNA-Nebularin, PNA-Inosine, PNA-4-Nitrobenzimidazole, PNA-3-Nitropyrrole, Morphorino-5-Nitroindole, Morphorino-Nebularin, Morphorino-Inosine, Morphorino-4-Nitrobenzimidazole, Morphorino-3-Nitropyrrole, Phosphoramidate-5-Nitroindole, Phosphoramidate-Nebularin, Phosphoramidate-Inosine, Phosphoramidate-4-Nitrobenzimidazole, Phosphoramidate-3-Nitropyrrole, 2'-O-MethoxyethylInosine, 2'-O-MethoxyethylNebularin, 2'-0-methoxyethyl 5-nitroindole, 2'-0-methoxyethyl 4-nitro-benzimidazole, 2'-0-methoxyethyl 3-nitropyrrole, and combinations of the above bases. More specifically, the above universal base is deoxyinosine, inosine, or a combination thereof.
[0043] Designing oligonucleotides used to detect specific target nucleic acid molecules can be carried out using various methods known in the art. For example, multiple target nucleic acid sequences for a specific target nucleic acid molecule can be collected and aligned, and then an oligonucleotide can be designed from each of the multiple target nucleic acid sequences to satisfy the design conditions. Therefore, to design an oligonucleotide, it is important to collect multiple target nucleic acid sequences for the target nucleic acid molecule of the organism of interest.
[0044] The oligonucleotide to be designed comprises a probe designed to satisfy at least one of the following conditions: (i) a Tm value at 50–85°C, (ii) a length of 15–50 nucleotides, and (iii) a mononucleotide (G). n Exclusion of run sequence; above, n is at least 3, (iv) the 5'-terminus is G or C, and (v) the GC content of the 5'-terminus is 40% or more.
[0045] More specifically, the probe design conditions include at least two of the conditions described above, more specifically at least three, even more specifically at least four, and even more specifically five conditions.
[0046] Among the above design conditions, the Tm value is, for example, 50-80℃, 50-75℃, 55-80℃, 55-75℃, 60-80℃, 60-75℃, 65-80℃, or 65-75℃. Specifically, among the design conditions, the Tm value is 55-80℃, 60-78℃, 63-78℃, 65-75℃, 67-75℃, or 65-73℃.
[0047] The length of the above design conditions is, for example, 10-60 nucleotides, 10-50 nucleotides, 10-45 nucleotides, 10-40 nucleotides, or 10-35 nucleotides, 15-60 nucleotides, 15-50 nucleotides, 15-45 nucleotides, 15-40 nucleotides, or 15-35 nucleotides.
[0048] Among the above design conditions, for example, a mononucleotide (G) where n is at least 3 or 4 n It is run sequence exclusion.
[0049] The GC content of the 5'-terminal region of the probe is 40% or more, specifically 40-70% or 40-60%. The 5'-terminal region refers to a region within 10 nucleotides from the 5'-terminal of the probe.
[0050] The oligonucleotide to be designed includes a primer designed to satisfy at least one of the following conditions: (i) a Tm value at 40–70°C, (ii) a length of 15–60 nucleotides, and (iii) a mononucleotide (G). n Exclusion of run sequence; above, n is at least 3.
[0051] Among the above design conditions, the Tm value is, for example, 40-70℃, 50-70℃, 55-70℃, 45-65℃, 50-65℃, 55-65℃, 45-60℃, or 50-65℃. Specifically, among the design conditions, the Tm value is 40-70℃, 45-65℃, 50-65℃, 50-60℃, 55-65℃, or 55-60℃.
[0052] The length of the above design conditions is, for example, 15-60 nucleotides, 15-50 nucleotides, 15-45 nucleotides, 15-40 nucleotides, 15-35 nucleotides, 15-30 nucleotides, 15-25 nucleotides, 18-45 nucleotides, 18-40 nucleotides, 18-35 nucleotides, 18-30 nucleotides, or 18-25 nucleotides. Specifically, the length of the design conditions is 15-40 nucleotides, 16-40 nucleotides, 17-40 nucleotides, 18-40 nucleotides, 15-35 nucleotides, 16-35 nucleotides, 17-35 nucleotides, 18-35 nucleotides, 15-30 nucleotides, 16-30 nucleotides, 17-30 nucleotides, 18-30 nucleotides, 18-25 nucleotides, or 17-25 nucleotides.
[0053] Mononucleotide (G) among the above design conditions n The criteria for the run sequence is, for example, a mononucleotide (G) where n is at least 3 or 4 n It is run sequence exclusion.
[0054] If the above primer is a DPO primer developed by the applicant (cf. U.S. Patent No. 8092997), the description of the Tm and length of the DPO primer disclosed in said patent document may be presented as the above design conditions.
[0055] More specifically, the primer design conditions include at least two of the conditions described above, and more specifically, at least three conditions.
[0056] In this specification, the term “sequence” refers to a specific arrangement of monomers within a macromolecule. In this specification, the term “nucleic acid sequence” refers to the arrangement of nucleotides within a nucleic acid molecule, representing the nucleic acid molecule as a specific nucleic acid sequence.
[0057] In this specification, the terms “nucleic acid sequence” or “nucleic acid sequence data” refer to the arrangement order of nucleotides within a nucleic acid molecule or information regarding the arrangement order of nucleotides within a nucleic acid molecule, and may be used interchangeably. The term “nucleic acid sequence data set” refers to a set of said nucleic acid sequence data, and said nucleic acid sequence data set may be provided in the form of a list of nucleic acid sequence data or an alignment file.
[0058] FIG. 3 is a flowchart of the processes for carrying out the method of the present invention according to one embodiment of the present invention. The method of the present invention will be described in detail with reference to FIG. 3 as follows:
[0059] Step (a): Collect synonyms for target nucleic acid molecules of the organism of interest ( 110 )
[0060] First, the method of the present invention receives the name of a target nucleic acid molecule and the name of an organism of interest, and retrieving synonyms of the organism of interest for the target nucleic acid molecule.
[0061] In this specification, “name of target nucleic acid molecule” refers to a word or symbol representing the target nucleic acid molecule. In the present invention, the name of the target nucleic acid molecule may be the name of a nucleotide molecule (e.g., a gene, a pseudogene, a non-coding sequence molecule, an untranscribed region, and a region of the genome). The name includes an official full name and a common name. The common name refers to a name used to represent the target nucleic acid molecule in the technical field to which the present invention belongs, other than the official name. In this specification, a symbol is a mark, sign, letter, or combination of letters representing the target nucleic acid molecule. The symbol includes an official symbol and an alias. An alias refers to an unofficial symbol used to identify the target nucleic acid molecule in the technical field to which the present invention belongs, other than the official symbol.
[0062] In this specification, the name of the organism of interest refers to the scientific name or taxonomic name of the organism according to a biological classification system. The name of the organism of interest also includes the taxonomic ID assigned to the scientific name or taxonomic name of the organism.
[0063] When implementing the computer-implemented method of the present invention, a user must input the name of the target nucleic acid molecule and the name of the organism of interest through a user interface (UI), and through such input, the computer implementing the method of the present invention receives the name of the target nucleic acid molecule and the name of the organism of interest through the user interface (UI).
[0064] Accordingly, in relation to the input of the name of a target nucleic acid molecule and the name of an organism of interest in this specification, receiving the name of the target nucleic acid molecule and the name of the organism of interest means receiving (or inputting into a computer) the name (word or mark) representing the target nucleic acid molecule and the name of the organism containing the target nucleic acid molecule. By such reception, it is determined whether to provide multiple target nucleic acid sequence data for a certain nucleic acid molecule by the method of the present invention.
[0065] In this specification, the terms “target nucleic acid sequence” or “target sequence” refer to a target nucleic acid molecule represented as a specific nucleic acid sequence.
[0066] A single target nucleic acid molecule, such as a single target gene, may have a single specific target nucleic acid sequence, or, in the case of a target nucleic acid molecule exhibiting genetic diversity or genetic variability, may have multiple diverse target nucleic acid sequences. The multiple target nucleic acid sequences in the present invention are target nucleic acid sequences having sequence similarity.
[0067] The method of receiving the name of the target nucleic acid molecule and the name of the organism containing the target nucleic acid molecule is not particularly limited; for example, it may be received by a user directly inputting it through an input device (e.g., UI), or it may be provided through various data storage media. Alternatively, the name of the target nucleic acid molecule and the name of the organism containing the target nucleic acid molecule may be received through wired or wireless data transmission.
[0068] Based on the name of the received target nucleic acid molecule and the name of the organism of interest, synonyms for the target nucleic acid molecule of the organism of interest are collected.
[0069] The method of the present invention is intended to collect and organize as much diverse target nucleic acid sequence data as possible so that oligonucleotides designed therefrom can have broad coverage of target nucleic acid molecules. Therefore, it is first necessary to collect as much target nucleic acid sequence data as possible, and to this end, synonyms for said target nucleic acid molecules to be used for collecting nucleic acid sequence data for oligonucleotide design are collected.
[0070] Synonyms for a target nucleic acid molecule refer to a group of words having the same meaning as a name or mark that identifies or designates the target nucleic acid molecule. In the present invention, synonyms for a target nucleic acid molecule refer to a group of words that includes both names and marks capable of identifying the target nucleic acid molecule, and include the official full name, common name, official symbol, and alias.
[0071] According to one embodiment of the present invention, synonyms for the target nucleic acid molecule are collected from a first database.
[0072] Synonyms for target nucleic acid molecules may be collected from a database. The database refers to a collection of organized data. The database may be a collection of organized data in which data is stored and accessible through a computer system. In this specification, to distinguish it from other databases, the database in which the synonyms are collected is referred to as the first database.
[0073] More specifically, the first database may be a genetic database.
[0074] A gene database refers to a database that collects, classifies, and stores information regarding genes contained in an organism. The gene database includes names, markers, and names of the organisms regarding the genes, and may include descriptions of the genes, information regarding the nucleic acid sequences of the genes (e.g., nucleic acid sequence identifiers), and information regarding proteins encoded by the genes (e.g., protein names, protein identifiers). The gene database may be referred to as a “database providing genetic information of an organism,” and the terms “gene database” and “database providing genetic information of an organism” may be used interchangeably in this specification.
[0075] According to one embodiment of the present invention, the first database may be a database that provides genetic information of an organism, including the title of a nucleic acid molecule, the name of a nucleic acid molecule, a description of a nucleic acid molecule, the name of an organism, and the name of a protein encoded by the nucleic acid molecule. The title is information that is listed as the title of a record when the first database provides information about a nucleic acid molecule to a user as a record.
[0076] The first database mentioned above may be a database built directly or a private genetic database with restricted users. Alternatively, the first database mentioned above may be public. The first database mentioned above includes not only those operated by the state or public institutions, but also those built by companies, educational institutions, research institutes, etc. According to one embodiment of the present invention, the first database mentioned above may be a publicly accessible genetic database selected from the group consisting of GenBank, EMBL, and DDBJ (DNA DataBank of Japan), or a genetic database built by downloading the publicly accessible genetic database.
[0077] According to one embodiment of the present invention, the first database may be a publicly accessible database containing the name of a target nucleic acid molecule and organism information, or a gene database constructed by downloading the same.
[0078] According to the present invention, synonyms for target nucleic acid molecules are retrieved from a first database, specifically automatically.
[0079] For example, a registrant registering the sequence of a nucleic acid molecule in a nucleotide database may list the official name or official mark in the field for listing the gene name of the nucleic acid molecule. However, in some cases, the official name of the nucleic acid molecule may be listed in a different field, and a different synonym of the nucleic acid molecule may be entered in the field for listing the gene name. Alternatively, the name of the target nucleic acid molecule received to execute the method of the present invention may not be the official name or official mark of the actual target nucleic acid molecule, but one of the other synonyms. Therefore, in order to obtain nucleic acid sequence data for a nucleic acid molecule, it is necessary to obtain as many synonyms as possible and use them for searching.
[0080] Therefore, (i) first, as many gene information summary records as possible related to the names of the input organism of interest and target nucleic acid molecules must be secured, and (ii) synonyms must be effectively secured from the secured gene information summary records.
[0081] First, to secure a sufficient number of genetic information summary records, we analyzed the top entries among the various input fields of genetic information summary records for nucleic acid molecules that had a high frequency of inputting the official name or official marker of the nucleic acid molecule. As a result, it was found that within the genetic database, entries containing the gene name of the nucleic acid molecule and entries containing the name of the protein produced by the expression of the nucleic acid molecule were the items with a high frequency of inputting the official name or official marker of the nucleic acid molecule. Specifically, it was confirmed that limiting the organism name to the organism of interest and collecting the results of searching for the name of the target nucleic acid molecule using the gene name, title, or protein name as the search field as genetic information summary records is the most effective method for collecting genetic information summary records for the target nucleic acid molecule.
[0082] Secondly, to identify synonyms from the acquired genetic information summary records, we analyzed the input items within the records that frequently contained names or markers other than the official names or markers of the target nucleic acid molecules. As a result, it was found that the entries containing the gene names of nucleic acid molecules and the description entries containing the names of proteins were the items with a high frequency of containing names or markers other than the official names or markers.
[0083] Therefore, it was confirmed that collecting information recorded in the gene names and descriptions of the collected gene information summary records is the most effective method.
[0084] According to one embodiment of the present invention, step (a) may include: (a-1) collecting a gene information summary record which is a record relating to the organism of interest and in which the name of the received target nucleic acid molecule is listed in the title, gene, or protein entry; and (a-2) collecting information listed in the name, symbol, and description of the gene information summary record to collect synonyms for the target nucleic acid molecule.
[0085] The above gene information summary record is a unit of edited information regarding a specific gene. According to one embodiment of the present invention, the gene information summary record is a unit of edited information regarding a specific gene that includes the name of the gene, the name of the protein, and descriptive information of the gene. The above gene information summary record is also referred to as a “gene information report” or a “gene report,” and “gene information summary record,” “gene information report,” or “gene report” may be used interchangeably in this specification.
[0086] Figure 4 shows the name of the target nucleic acid molecule (ompA) and the name of the organism of interest ( Chlamydophila pneumoniae As a result of inputting ), gene information summary records are collected from the gene database of NCBI (National Center for Biotechnology Information) or from a gene database built by downloading the said gene database, and protein names are collected from gene descriptions as synonyms for target nucleic acid molecules in said records.
[0087] According to one embodiment of the present invention, step (a) may be a step in which a computer processor communicates with a first database via a wired or wireless network to collect synonyms based on the name of a specified target nucleic acid molecule and the name of an organism of interest.
[0088] The above reception can be performed by receiving the name of the target nucleic acid molecule and information regarding the organism of interest directly from the user, or by receiving it in the form of a file.
[0089] According to one embodiment of the present invention, step (a) may include the following steps: sending a command to the first database to send a gene information summary record to the memory of a computer, wherein the record is a record relating to an organism of interest received among the gene records included in the first database, and at least one of the title, gene, and protein items of the record is identical to the name of the received target nucleic acid molecule; receiving the gene information summary record sent by the first database in response to the command; and collecting information described in the name, symbol, and description items of the received gene information summary record and storing it in memory as synonyms.
[0090] The above transmission and reception can be performed via wired or wireless networks.
[0091] According to one embodiment of the present invention, the collection of synonyms in step (a) may be carried out by transmitting a command to a first database to transmit to a computer memory information described in the name, symbol, and description items of a gene information summary record, which is a record relating to a received organism of interest among the gene records included in a first database, and in which at least one of the title, gene, and protein items of the record is identical to the name of the received target nucleic acid molecule, and collecting the information sent by the first database in response to the command and storing synonyms in memory. The synonyms collected in memory may be stored in a storage medium in the form of electronic files.
[0092] According to one embodiment of the present invention, step (a) may be a step of receiving the name of a target nucleic acid molecule, the name of a protein encoded by the target nucleic acid molecule, and the name of a source organism, and retrieving synonyms of the organism for the target nucleic acid molecule and the protein.
[0093] According to the present embodiment, in addition to the name of the target nucleic acid molecule and the name of the organism, the user may input the name of the protein encoded by the target nucleic acid molecule that is known to the user.
[0094] According to the present embodiment, the collection of synonyms in step (a) may be carried out by sending a command to the first database to transmit to the memory of a computer information described in the name, symbol, and description items of a gene information summary record, which is a record concerning a received organism of interest among the gene records included in the first database, and in which at least one of the title, gene, and protein items of the record is identical to the name and protein name of the received target nucleic acid molecule, and in response to the command, collecting the information sent by the first database and storing synonyms in the memory.
[0095] The user can review the synonyms of the names of target nucleic acid molecules and / or proteins collected in this manner and delete any synonyms deemed inappropriate so that they are not considered in the subsequent steps.
[0096] Step (b): Collect nucleic acid sequence data included in nucleic acid records ( 120 )
[0097] Next, the method of the present invention comprises the step of (b) collecting nucleic acid sequence data contained in nucleic acid records. Each of the nucleic acid records relates to the organism of interest and includes a descriptor in which the name of the target nucleic acid molecule and at least one of the collected synonyms are listed.
[0098] According to the present invention, the nucleic acid sequence data collected in step (b) is used only to select taxonomic representative sequences and group representative sequences as described below, and is not provided as a target nucleic acid sequence data set for target nucleic acid molecules used in the design of oligonucleotides.
[0099] According to one embodiment of the present invention, the nucleic acid sequence data of step (b) includes nucleic acid sequence data corresponding to part or all of the target nucleic acid molecule or variant nucleic acid sequence data for the target nucleic acid molecule.
[0100] The variant nucleic acid sequence data for the above target nucleic acid molecule represents nucleic acid sequence data comprising a nucleotide sequence in which one or more nucleotides are substituted, deleted, and / or added compared to the target nucleic acid sequence of the target nucleic acid molecule.
[0101] According to one embodiment of the present invention, nucleic acid sequence data included in the nucleic acid records are collected from a second database.
[0102] The second database, which is distinguished from the first database described above, refers to a nucleotide database that collects, classifies, and stores nucleic acid sequence data of various nucleic acid molecules, and the second database may be used interchangeably with “nucleotide database,” “nucleic acid sequence database,” or “nucleic acid information collection” in this specification.
[0103] The second database described above includes nucleic acid records. The nucleic acid records may include nucleic acid sequence data for a nucleic acid molecule and metadata regarding the nucleic acid sequence data as a descriptor. Metadata regarding the nucleic acid sequence data as a descriptor refers to bibliographic information regarding the nucleic acid sequence data, and the metadata may include, for example, an identifier for the nucleic acid sequence data, information about an organism containing the corresponding nucleic acid molecule, keywords, and information regarding references such as papers in which the nucleic acid sequence data has been published.
[0104] The above nucleic acid record may be referred to as a “nucleic acid report” or a “nucleic acid information report,” and in this specification, “nucleic acid record,” “nucleic acid report,” or “nucleic acid information report” may be used interchangeably.
[0105] In this specification, a descriptor refers to an item that describes or identifies specific nucleic acid sequence data. The descriptor is metadata for specific nucleic acid sequence data and, specifically, may be any field of a nucleic acid record containing specific nucleic acid sequence data. More specifically, the descriptor may include a name, definition, keywords, source organism, and reference title.
[0106] The second database may be a database built directly or a private nucleotide database with restricted users. Alternatively, the second database may be public. The public second database includes not only those operated by the state or public institutions, but also those built by companies, educational institutions, research institutes, etc. According to one embodiment of the present invention, the second database may be a publicly accessible nucleotide database selected from the group consisting of GenBank, EMBL, and DDBJ, or a nucleotide database built by downloading the publicly accessible nucleotide database. According to another embodiment of the present invention, the second database may be a nucleotide database including NCBI's GenBank (including STS, EST, GSS, SNP, TSA, PAT, WGS, and non-WGS databases), RefSeq, DDBJ, and EMBL databases, or a nucleotide database built by downloading the nucleotide database.
[0107] In the present invention, the first database and the second database may be from the same institution, or databases provided by different institutions may be used.
[0108] According to one embodiment of the present invention, the collection of nucleic acid sequence data of step (b) is carried out by a method comprising the following steps:
[0109] (b-1) a step of collecting identifiers of nucleic acid records; each of the nucleic acid records relates to the organism of interest and includes a descriptor in which the name of the target nucleic acid molecule and at least one of the collected synonyms are listed, and
[0110] (b-2) A step of collecting nucleic acid sequence data identified by the above-mentioned identifiers.
[0111] Step (b-1): Collect the identifiers of nucleic acid records
[0112] According to one embodiment of the present invention, the identifiers of the nucleic acid records are collected from a second database.
[0113] In the present invention, an identifier is data used to identify specific nucleic acid sequence data. The data used as the identifier has no special restrictions on its format, such as characters, numbers, or combinations thereof. Different identifiers may be assigned to the same nucleic acid sequence data depending on the database. Examples of the identifier include an accession number, an accession version, or a GI number.
[0114] According to one embodiment of the present invention, the second database may be a database providing nucleic acid records including nucleic acid sequence data and an identifier and a descriptor for said nucleic acid sequence data.
[0115] Figure 5 shows the organism of interest Enterobacter cloacae complexIf the target nucleic acid molecule is ompX, it shows the collection of identifiers of nucleic acid records from a second database. Looking at Figure 5, Accession No and GI No can be identified as identifiers.
[0116] Figure 6 is a captured image of a portion of the nucleic acid record that appears when the title of Accession: CP017990.1 in Figure 5 is clicked. In the above nucleic acid record gene , CDS, / gene, / note, / product, etc. represent descriptors.
[0117] According to one embodiment of the present invention, step (b-1) may be a step in which a computer processor communicates with a second database via a wired or wireless network based on information regarding the name of the collected target nucleic acid molecule and synonyms thereof and the received organism, and collects identifiers of nucleic acid records from the second database in which at least one of the name of the target nucleic acid molecule and the collected synonyms is listed in a descriptor.
[0118] According to one embodiment of the present invention, step (b-1) may be a step in which a process transmits a command to a second database to transmit an identifier of a nucleic acid record satisfying the following conditions to the memory of a computer, and receives information sent by the second database in response to the command via a wired or wireless network and stores it in memory: (i) that the received record is of an organism of interest; (ii) that at least one of the name of the received target nucleic acid molecule or the collected synonyms is a nucleic acid record listed in a descriptor (i.e., metadata).
[0119] The identifiers collected in the above memory can be stored in a storage medium in the form of electronic files.
[0120] According to one embodiment of the present invention, step (b-1) relates to the organism of interest and is a step of collecting identifiers of nucleic acid records in which at least one of the name of the target nucleic acid molecule, the name of the protein, and the collected synonyms is listed in the descriptor.
[0121] Step (b-2): Collect nucleic acid sequence data identified by identifiers
[0122] The above-mentioned identifiers are the identifiers collected in step (b-1) and are identifiers that indicate target nucleic acid sequence data for the target nucleic acid molecule. Therefore, the nucleic acid sequence data specified by the identifiers in step (b-2) are target nucleic acid sequence data for the target nucleic acid molecule.
[0123] In International Publication No. WO2019 / 212238 filed by the present applicant, nucleic acid sequence data specified by the above-mentioned identifiers are collected and provided as a target nucleic acid sequence data set for a target nucleic acid molecule for use in the design of an oligonucleotide; however, according to the present invention, the nucleic acid sequence data collected in step (b-2) is used only to select taxonomic representative sequences and group representative sequences as described below, and is not provided as a target nucleic acid sequence data set for a target nucleic acid molecule for use in the design of an oligonucleotide.
[0124] The collected nucleic acid sequence data above may be all or part of the nucleic acid sequence data specific to the identifiers.
[0125] The collection of nucleic acid sequence data specified by the above-mentioned identifiers is carried out from a second database using the identifiers.
[0126] The above nucleic acid sequence data may be collected by collecting the nucleic acid sequence itself, which is identified by identifiers from the second database, or by collecting nucleic acid records identified by identifiers from the second database and extracting the nucleic acid sequence therefrom.
[0127] According to one embodiment of the present invention, step (b-2) may involve the process sending a command to a second database requesting nucleic acid sequence data specified by the identifiers, and receiving the nucleic acid sequence data sent by the second database in response to the command and storing it in memory.
[0128] According to one embodiment of the present invention, step (b-2) selectively collects nucleic acid sequence data corresponding to the target nucleic acid molecule within the nucleic acid sequence data specified by the identifiers.
[0129] The nucleic acid records collected by the above-mentioned identifiers may contain only nucleic acid sequence data corresponding to the target gene, which is the target nucleic acid molecule; however, in many cases, they may contain nucleic acid sequence data of other genes in addition to the sequence of the target gene. Furthermore, depending on the purpose of collecting the target nucleic acid sequence data set for the target nucleic acid molecule, there may be cases where only nucleic acid sequence data corresponding to a part of the target nucleic acid molecule needs to be collected. Therefore, after collecting all or part of the nucleic acid records specified by the above-mentioned identifiers, nucleic acid sequence data corresponding to the target nucleic acid molecule suitable for the purpose can be selected and collected from therefrom.
[0130] According to one embodiment of the present invention, step (b-2) may include the following steps: (b-2-1) collecting nucleic acid records specified by the identifiers; and (b-2-2) collecting nucleic acid sequence data corresponding to a target nucleic acid molecule from the nucleic acid records.
[0131] According to one embodiment of the present invention, step (b-2-2) may selectively collect nucleic acid sequence data corresponding to a target nucleic acid molecule and identification information of the nucleic acid sequence data from the nucleic acid records.
[0132] In this specification, the expression “selectively collected” means collecting only the necessary nucleic acid sequence data among the nucleic acid sequence data within the nucleic acid record.
[0133] When the nucleic acid sequence data included within a nucleic acid record comprises multiple distinguishable nucleic acid sequence data, the nucleic acid record comprises multiple sub-records. In this specification, the term “sub-record” refers to a unit of data group comprising distinguishable nucleic acid sequence data and / or details thereof within a single nucleic acid record. Each sub-record includes location information regarding the nucleic acid sequence data corresponding to that sub-record and details describing the nucleic acid sequence data corresponding to that sub-record.
[0134] In this specification, “distinguishable nucleic acid sequence data” refers to each of the two or more nucleic acid sequences and their detailed items when two or more nucleic acid sequences that can be recognized as physically or functionally different are included within a single nucleic acid record. For example, when nucleic acid sequences for multiple genes encoding different proteins are all included in a single nucleic acid sequence data within a nucleic acid record, the nucleic acid sequence data can be distinguished into parts corresponding to each gene.
[0135] In order to selectively collect target nucleic acid sequence data from a single nucleic acid record containing multiple distinguishable nucleic acid sequence data, it is determined whether a corresponding sub-record is a valid sub-record based on the descriptors (specifically, details) included in each sub-record within the nucleic acid record. In this specification, "valid sub-record" refers to a sub-record that contains the nucleic acid sequence data to be collected or contains location information regarding it.
[0136] The descriptor (specifically, the sub-item) included in each sub-record refers to an item in which information regarding nucleic acid sequence data containing location information is recorded within each sub-record. The above sub-item may include, for example, the gene name indicated by the sub-record, information on the protein produced from the gene indicated by the sub-record (e.g., protein name, protein identifier), records of the nucleic acid record provider, amino acid sequence information, etc.
[0137] The inventors confirmed that the collected synonyms are listed in some of the detailed items of the sub-record, and also confirmed that the frequency and accuracy of the listing of the collected synonyms vary by detailed item. Therefore, they discovered that the most efficient method for obtaining the intended nucleic acid sequence data is to select some of the detailed items, assign priority to them, and sequentially check whether the collected synonyms are listed in those detailed items.
[0138] The inventors compared the genes included in the nucleic acid sequence data indicated by each sub-record with the data of the sub-items of the corresponding sub-record and confirmed that the collected synonyms are listed most frequently in the sub-items regarding gene names, followed by the sub-items regarding protein information, and then the sub-items regarding records of nucleic acid record providers. Therefore, they determined that checking whether the collected synonyms are listed in the sub-items in this order to determine valid sub-records, and collecting nucleic acid sequence data and identification information for the determined valid sub-records, is the most efficient method for selectively obtaining the desired nucleic acid sequence data accurately.
[0139] Accordingly, according to one embodiment of the present invention, the step of selectively collecting nucleic acid sequence data corresponding to a target nucleic acid molecule and identification information of said nucleic acid sequence data from said nucleic acid records may include the following steps:
[0140] (b-2-2-1) A step of determining, among one or more sub-records within the nucleic acid record, a sub-record in which the synonym is recorded in a predetermined first sub-item as a valid sub-record;
[0141] (b-2-2-2) If there is no valid subrecord determined by the first sub-item within the nucleic acid record, a step of determining a subrecord in which the synonym is recorded in the second sub-item as a valid subrecord;
[0142] (b-2-2-3) If there is no valid subrecord determined by the second sub-item within the nucleic acid record, a step of determining a subrecord in which the synonym is recorded in the third sub-item as a valid subrecord; and
[0143] (b-2-2-4) A step of collecting nucleic acid sequence data and identification information for one valid subrecord determined above.
[0144] According to another embodiment of the present invention, the step of selectively collecting nucleic acid sequence data corresponding to a target nucleic acid molecule and identification information of said nucleic acid sequence data from said nucleic acid records may include the following steps:
[0145] (b-2-2-1) A step of determining, among one or more sub-records within the nucleic acid record, a sub-record in which the synonym is recorded in a predetermined first sub-item and a second sub-item as a valid sub-record;
[0146] (b-2-2-2) If there is no valid subrecord determined by the first and second subrecords within the nucleic acid record, a step of determining a subrecord in which the synonym is recorded in the third subrecord as a valid subrecord; and
[0147] (b-2-2-3) A step of collecting nucleic acid sequence data and identification information for one valid subrecord determined above.
[0148] According to one embodiment of the present invention, the first item is a item related to the name of a nucleic acid sequence within a subrecord, the second item is a item related to protein information produced from said gene, and the third item may be a item related to a note of a genetic information provider.
[0149] According to a more embodiment of the present invention, the step of selectively collecting nucleic acid sequence data corresponding to a target nucleic acid molecule and identification information of said nucleic acid sequence data from said nucleic acid records may include the following steps:
[0150] (b-2-2-1) A step of determining as valid subrecords, among one or more subrecords within the nucleic acid record, a subrecord in which the name of a target nucleic acid molecule and / or a synonym thereof is recorded in a first sub-item related to a predetermined gene name, and a subrecord in which the name of a protein of the target nucleic acid molecule and / or a synonym thereof is recorded in a second sub-item related to protein information;
[0151] (b-2-2-2) If there is no valid subrecord determined by the first and second subrecords within the nucleic acid record, a step of determining as a valid subrecord a subrecord in which at least one of the target nucleic acid molecule name, the protein name of the target nucleic acid molecule, and a synonym thereof is recorded in the third subrecord related to the genetic information provider's note; and
[0152] (b-2-2-3) A step of collecting nucleic acid sequence data and identification information for one valid subrecord determined above.
[0153] According to the present embodiment, if the gene name listed in the first sub-item related to the gene name among the descriptors in the nucleic acid record matches the name of the target nucleic acid molecule received in step (a) and / or a synonym of the target nucleic acid molecule collected in step (a), and the protein name listed in the second sub-item related to the protein information matches the name of the protein received in step (a) and / or a synonym of the protein name collected in step (a), then nucleic acid sequence data corresponding to the target nucleic acid molecule and identification information of the nucleic acid sequence data can be selectively collected from the nucleic acid records.
[0154] As a result, when collecting nucleic acid sequence data in which at least one of the name of the target nucleic acid molecule and collected synonyms is listed in the gene name, protein name, and the genetic information provider's note among the descriptors within the nucleic acid record, it is possible to collect a target nucleic acid sequence more suitable for oligonucleotide design.
[0155] According to one embodiment of the present invention, the method may additionally include, after step (b-2-2), a step (b-2-2A) of merging nucleic acid sequence data in which part or all of the nucleic acid sequence data among the plurality of nucleic acid sequence data overlaps to provide nucleic acid sequence data, where the nucleic acid sequence data collected from one nucleic acid record of step (b-2-2) is a plurality of nucleic acid sequence data.
[0156] For example, the existence of multiple nucleic acid sequence data regarding a single target nucleic acid molecule within a single nucleic acid record corresponds to a case where each of the nucleic acid sequence data is a partial sequence of said target nucleic acid molecule. In other words, rather than the multiple nucleic acid sequence data each independently encoding said target nucleic acid molecule-encoding protein, the entirety of the multiple nucleic acid sequence data encodes a single target nucleic acid molecule-encoding protein. In such a case, it is not appropriate to use each of the multiple nucleic acid sequence data as the target nucleic acid sequence for the target nucleic acid molecule, and it is preferable to treat a single nucleic acid sequence formed by merging them as the target nucleic acid sequence for the target nucleic acid molecule. For example, regarding the same gene or descriptor gene The location information listed in is 1 to 10, and the descriptor is CDSAccording to the present invention, when the position information described herein is 2 to 8, both the nucleic acid sequence data corresponding to position information 1 to 10 and the nucleic acid sequence data corresponding to position information 2 to 8 can be collected. When such two nucleic acid sequence data are provided as a nucleic acid sequence data set for oligonucleotide design, a problem arises in which, even though there is only one nucleic acid sequence having the same identifier, two nucleic acid sequence data are provided, and for example, the oligonucleotide designed at position information 2 to 8 has a target coverage of 2 instead of 1. To resolve this problem, in the case of nucleic acid sequences having the same identifier, the overlapping parts are merged to provide only one nucleic acid sequence data corresponding to position information 1 to 10.
[0157] In this specification, the term “target-coverage” means the ratio of a plurality of nucleic acid sequences to a target nucleic acid molecule that matches the sequence of an oligonucleotide (specifically, 100% match, 95% or more match, 90% or more match, etc.).
[0158] According to a more specific embodiment of the present invention, the merging of step (b-2-2A) is carried out by a method comprising the following steps:
[0159] (b-2-2A-1) A step of collecting sequence position information (specifically, position information of the start point and end point) of each of the plurality of nucleic acid sequence data corresponding to the target nucleic acid molecule within the above-mentioned single nucleic acid record;
[0160] (b-2-2A-2) A step of analyzing the sequence position information above to select nucleic acid sequence data among the plurality of nucleic acid sequence data in which some or all of the sequence data overlap each other; and
[0161] (b-2-2A-3) A step of generating new nucleic acid sequence data that includes all of the selected nucleic acid sequence data.
[0162] As described above, by providing a single nucleic acid sequence data that merges nucleic acid sequence data that partially or wholly overlaps, target nucleic acid sequence data for parts of a target nucleic acid molecule are recognized as target nucleic acid sequence data for independent target nucleic acid molecules, thereby reducing the probability of statistical error occurring in analysis based on the target nucleic acid sequence data set.
[0163] According to one embodiment of the present invention, between steps (b) and (c), the following step is additionally included:
[0164] (b-3) A step of sorting the collected nucleic acid sequence data according to biosample identifiers to select nucleic acid sequence data having the same biosample identifier;
[0165] (b-4) A step of aligning the selected nucleic acid sequence data to satisfy at least one of the following alignment criteria;
[0166] (b-5) A step of selecting the nucleic acid sequence of the highest nucleic acid sequence data among the above-mentioned aligned nucleic acid sequence data; and
[0167] (b-6) A step of removing nucleic acid sequence data other than the top-most nucleic acid sequence data from the collected nucleic acid sequence data, and the alignment criteria include the following:
[0168] (i) sort the selected nucleic acid sequence data according to the assembly level; the assembly levels are ranked in the order of complete genome, chromosome, scaffold, and contig (i.e., the complete genome has the highest rank),
[0169] (ii) Sort the selected nucleic acid sequence data based on whether the selected nucleic acid sequence data is included in the Reference Sequence (RefSeq) database; if the nucleic acid sequence data is included in the Reference Sequence (RefSeq) database, the ranking is higher than if it is not included.
[0170] Considering the characteristics of the genome sequencing process according to the genome project, registration in a database, and transfer to another database after verification of the registered sequence information, the inventors have confirmed that even if the access number for each nucleic acid sequence is different, if the information for the biosample is the same, the assembly level of the genomic sequence differs depending on the progress of genome sequencing and the database in which it is stored is different. By utilizing these characteristics, they intend to remove nucleic acid sequence data other than the nucleic acid sequence that has reached the most recent genome sequencing progress from the nucleic acid sequence data collected in step (b).
[0171] Specifically, when a person who has identified gene and protein information, etc. of a genomic sequence under a genome project intends to register information regarding a nucleotide sequence in a database, they input a bioproject and a biosample regarding the collected information of the said genomic sequence, and said bioproject and biosample have a unique number consisting of a combination of letters and numbers.
[0172] Furthermore, during the whole genome sequencing process under the genome project, genomic sequences are fragmented into sub-sequences to determine whether they are proteins-coding genes and what functions those proteins perform. Afterward, these fragmented sequences are merged. In this process, nucleic acid sequences acquire assembly levels in the order of contig, scaffold, chromosome, and complete genome. In other words, the order of these assembly levels indicates the progress of the genome sequencing process (the complete genome has the highest assembly level).
[0173] In addition, when the genome sequencing process of genomic sequences proceeds and results in chromosome and complete genome assembly levels, the data undergoes a transfer process from the initially registered database (e.g., NCBI’s GenBank) to another database (e.g., NCBI’s RefSeq), and nucleic acid sequence data at the assembly sub-level is deleted from the initially registered database.
[0174] However, even if transferred to NCBI's RefSeq database, it may still exist in NCBI's GenBank without being deleted.
[0175] And, these processes for genomic sequences exist as information about nucleic acid sequence data in descriptors within nucleic acid records, and the inventors have used this information to remove duplicate sequences in this embodiment.
[0176] For example, if the second database used in the present invention includes the NCBI GenBank and RefSeq databases, and even though the sequencing process for the genomic sequence of the organism of interest is completed and the nucleic acid sequence data of the complete genome is in the RefSeq database, but nucleic acid records at the assembly level of the contig, scaffold, and chromosome of the organism of interest still exist in the NCBI GenBank database, then the nucleic acid sequence data collected according to the present invention includes nucleic acid sequence data having the four types of assembly levels.
[0177] If nucleic acid sequence data having the four types of assembly levels mentioned above is provided for the design of an oligonucleotide, a problem arises in which the target coverage increases fourfold even though it is the same nucleic acid sequence, and if sequence information is modified at the chromosome level during the genome sequencing process and the modified sequence information is reflected at the complete genome level, and if an oligonucleotide is designed at the part where the sequence information is modified, problems may arise such as having to introduce a degenerate base at the location where the sequence information was modified.
[0178] In this case, duplicate sequences can be removed so as not to cause the aforementioned problems by utilizing the fact that the nucleic acid sequence data corresponding to the four types of assembly levels are nucleic acid sequence data for the same organism that has information about the same biosample.
[0179] The biosample identifier in step (b-3) above represents a unique biosample number, and if the nucleic acid sequence data in (ii) above is included in the RefSeq (Reference Sequence) database, a unique number is assigned to the nucleic acid record.
[0180] By removing redundant sequences according to this embodiment, the aspects of target coverage and the time and cost associated with introducing degenerate bases can be improved in designing oligonucleotides.
[0181] Step (c): Select taxonomic representative sequence ( 130 )
[0182] And, the method of the present invention (c) sorts the collected nucleic acid sequence data according to a taxonomic name and / or a taxonomic ID and selects a taxonomic representative sequence from among the nucleic acid sequence data having the same taxonomic name and / or taxonomic ID.
[0183] In this specification, the term “Taxonomic name” refers to the scientific name of an organism classified according to a biological classification system, and the term “Taxonomic ID” refers to the number assigned to the said taxonomic name. The taxonomic ID is linked, for example, with the organism information of nucleic acid records retrieved from NCBI and the NCBI Taxonomy database, so it can be verified through the Taxonomy viewer, and can also be verified in the taxon entry of the nucleic acid records retrieved from NCBI.
[0184] The collected nucleic acid sequence data are sorted according to taxonomic names and / or taxonomic IDs.
[0185] Then, among the nucleic acid sequence data aligned by the above taxonomic names and / or taxonomic identifiers, nucleic acid sequence data having the same taxonomic name and / or taxonomic identifier are classified, and a taxonomic representative sequence is selected from among them.
[0186] According to one embodiment of the present invention, the selection of the taxonomic representative sequence of step (c) is carried out by a method comprising the following steps:
[0187] (c-1) A step of aligning nucleic acid sequence data having the same taxonomic name and / or taxonomic ID to satisfy at least one of the following predetermined alignment criteria; and
[0188] (c-2) A step of selecting the nucleic acid sequence of the top-most nucleic acid sequence data among the above-mentioned aligned nucleic acid sequence data as a taxonomic representative sequence, and the above-mentioned predetermined alignment criteria include the following:
[0189] (i) sorted according to the assembly level of the above nucleic acid sequence data; the assembly levels are ranked in the order of complete genome, chromosome, scaffold, and contig (i.e., the complete genome has the highest rank),
[0190] (ii) Sort according to whether the above nucleic acid sequence data is included in the RefSeq (Reference Sequence) database; if the above nucleic acid sequence data is included in the RefSeq database, it is ranked higher than if it is not included, and
[0191] (iii) Sort according to whether the name of the nucleic acid molecule listed in the descriptor of the nucleic acid record containing the nucleic acid sequence data matches at least one of the name of the received target nucleic acid molecule and the collected synonyms; the case of a match is ranked higher than the case of a non-match, and
[0192] (iv) sorted according to the length of the nucleic acid sequence data; the longer the length, the higher the rank, and
[0193] (v) Sort according to whether a host is listed in the descriptor of the nucleic acid record containing the above nucleic acid sequence data; the case where the host of interest for the organism of interest is listed in the host is higher than the case where it is not listed, and the case where it is not listed is higher than the case where an organism other than the source organism of interest is listed in the host, and
[0194] (vi) Sort according to the registration date or modification date of the nucleic acid record containing the above nucleic acid sequence data; the newer the registration date or modification date, the higher the rank, and
[0195] (vii) Sort according to the alphabet of the access number of the above nucleic acid sequence data; the earlier the alphabet of the access number, the higher the rank.
[0196] According to the present embodiment, nucleic acid sequence data having the same taxonomic name and / or taxonomic ID are aligned to satisfy at least one alignment criterion (specifically, alignment criterion (i)), specifically at least two, more specifically at least three, at least four, at least five or at least six, most specifically at seven ranks.
[0197] According to one embodiment of the present invention, the at least two alignment criteria differ in criticality, and the method of the present invention additionally includes the step of aligning nucleic acid sequence data to satisfy the at least two alignment criteria considering the criticality.
[0198] There are two main methods for aligning nucleic acid sequence data with the same taxonomic name and / or taxonomic ID to select taxonomic representative sequences:
[0199] According to the first method, the above at least two alignment criteria differ in criticality, and nucleic acid sequence data can be aligned to satisfy the alignment criterion with the highest criticality (e.g., alignment criterion (i)).
[0200] If there are multiple nucleic acid sequence data that satisfy the alignment criterion of highest importance, the nucleic acid sequence data can be aligned to satisfy the next-highest alignment criterion.
[0201] For example, if the importance of alignment criteria is in the order of (i), (ii), (iii), (iv), (v), (vi), and (vii), and there are 3 nucleic acid sequence data that satisfy alignment criterion (i), these 3 nucleic acid sequence data are aligned according to alignment criterion (ii). If there are 3 nucleic acid sequence data that satisfy alignment criterion (ii), the nucleic acid sequence data are aligned to satisfy alignment criterion (iii).
[0202] According to the second method, by assigning different weights to alignment criteria, assigning scores to values (or ranges of values) within each alignment criterion, and considering the rankings thereof, a total score for each nucleic acid sequence data can be obtained; by considering this calculated total score, the nucleic acid sequence data can be aligned, and the top-ranked nucleic acid sequence data can be selected as the taxonomic representative sequence based on the ranking according to the total score.
[0203] Figure 7 shows the organism of interest ( Enterobacter cloacae This document shows a process of collecting nucleic acid sequence data identified by collected identifiers using the name (ompX) and synonyms (outer membrane protein) of the target nucleic acid molecule of the complex, aligning the nucleic acid sequence data according to the alignment criteria among the nucleic acid sequence data having the same taxonomic identifier, and selecting a taxonomic representative sequence according to the first method described above.
[0204] As shown in Fig. 7, the collected nucleic acid sequence data have the same taxonomic ID (Taxid) of 550. The alignment criteria have importance in the order of (i) to (vii). First, sort according to the assembly level of the nucleic acid sequence of alignment criterion (i) (the higher the assembly level, the higher the rank), sort according to whether it is included in the RefSeq (Reference Sequence) database of alignment criterion (ii) (if included in the RefSeq database, it has a unique number), sort according to whether the name of the nucleic acid molecule listed in the descriptor of the nucleic acid record of alignment criterion (iii) matches at least one of the name of the received target nucleic acid molecule and the collected synonyms (ranks are higher in the order that the name of the nucleic acid molecule listed in the descriptor of the nucleic acid record matches the protein name, the protein name matches, and the name of the nucleic acid molecule matches), sort according to the length of the nucleic acid sequence of alignment criterion (iv) (the longer the length, the higher the rank), and sort according to whether the host is listed in the descriptor of the nucleic acid record of alignment criterion (v) (if Homo sapiens is listed in the host, if an organism is not listed in the host Rankings are higher in the order of cases where no host is listed, and cases where an organism other than Homo sapiens is listed), sorted according to the registration or modification date of the nucleic acid records in alignment criterion (vi) (the newer the date, the higher the rank), and sorted according to the alphabet of the access number of the nucleic acid sequence in alignment criterion (vii) (the earlier the alphabet, the higher the rank).
[0205] As a result, the nucleic acid sequence data with access number CP040827.1 appeared to be at the top, and thus, this top-ranking nucleic acid sequence data is selected as the taxonomic representative sequence.
[0206] If the nucleic acid sequence data collected in step (b) above includes only nucleic acid sequence data having the same scientific name (same taxonomic name or taxonomic identifier) as the organism of interest, one taxonomic representative sequence may be selected.
[0207] According to one embodiment of the present invention, one or more taxonomic representative sequences may be selected from the collected nucleic acid sequence data.
[0208] If the nucleic acid sequence data collected in step (b) above includes nucleic acid sequence data of organisms having the same scientific name (same taxonomic name or taxonomic identifier) as the organism of interest, as well as organisms that are superordinate or subordinate to the organism of interest in the biological classification system, or organisms having a different scientific name (taxonomic name or taxonomic identifier) from the organism of interest, multiple taxonomic representative sequences may be selected. That is, a taxonomic representative sequence is selected for each scientific name (taxonomic name or taxonomic identifier) of the organism.
[0209] According to one embodiment of the present invention, the taxonomic representative sequence satisfies the following predetermined length criteria: within a predetermined range of the intermediate value of the nucleic acid sequence length having the assembly level of the complete genome and / or chromosome among the nucleic acid sequence data collected in step (b). The predetermined range is not particularly limited, but may be, for example, ± 2%, 4%, 5%, 10%, 15%, 20%, 25%, or 30% (bp, mer, or nucleotide length) of the intermediate value.
[0210] Step (d): Selecting the group representative ranking ( 140 )
[0211] Next, the method of the present invention (d) groups the selected taxonomic representative sequences according to homology and selects a group representative sequence from each group.
[0212] Since the collection of synonyms and the nucleic acid sequence data obtained based thereon are searched based on names, there is an advantage in that sequences with low mutual agreement due to significant variation between nucleic acid sequences can be collected. However, sequences registered with names that are not the names of known target nucleic acid molecules or synonyms, such as when the name of the nucleic acid molecule is not confirmed at the time of recording in the second database (specifically, the nucleotide database), or due to errors by the registrant, may not be collected even if the corresponding nucleic acid sequence is actually the target nucleic acid sequence for the target nucleic acid molecule.
[0213] To reinforce this aspect, a conventional method (WO2019 / 212238) collects nucleic acid sequence data based on synonyms for the target nucleic acid molecule of the organism of interest, aligns the collected nucleic acid sequence data along sequence lengths, determines the nucleic acid sequence with the longest sequence length as a representative sequence, groups the representative sequence and the collected nucleic acid sequence data according to homology, determines the nucleic acid sequence data with the longest sequence in each group as a group representative sequence, supplements the nucleic acid sequence data for each group by adding nucleic acid sequence data having homology greater than a predetermined value with the determined group representative sequence, and then provides a set of nucleic acid sequence data used for the design of oligonucleotides by adding the nucleic acid sequence data supplemented through the group representative sequence to the nucleic acid sequence data collected based on the synonyms.
[0214] As a result of providing the nucleic acid sequence data set provided through this process in the form of an alignment file, as can be seen in Figure 2, the alignment results of multiple nucleic acid sequence data were not properly formed, and in order to use this for designing oligonucleotides, analysts had to review the alignment results and check for errors such as the registration of group representative sequences.
[0215] Furthermore, the inventors reviewed the alignment results and confirmed that the alignment of the collected nucleic acid sequence data was not properly formed due to the group representative sequence selected by the aforementioned conventional method.
[0216] Accordingly, the inventors aligned sequences collected as synonyms according to taxonomic names and / or taxonomic identifiers, selected nucleic acid sequence data having the same taxonomic name and / or taxonomic identifier as taxonomic representative sequences, grouped the said taxonomic representative sequences according to homology, and selected group representative sequences therefrom, thereby confirming that the alignment result of the collected nucleic acid sequence data was properly formed.
[0217] According to the present invention, in order to select a group representative sequence, the selected taxonomic representative sequence is grouped according to homology.
[0218] In this specification, the term “homology” refers to a state in which two or more nucleic acid sequences are identical or similar in relative, positional, or structural terms. Homology can be expressed numerically as a degree of similarity or correspondence between two nucleic acid sequences, specifically as a ratio (percentage).
[0219] Specifically, the above homology may be identity or similarity. In this specification, the term “identity” is determined by whether the bases at specific positions of the two sequences being compared are identical. In this specification, the term “similarity” is determined by considering the characteristics of the bases at specific positions of the two sequences being compared to distinguish whether they are identical, have different but similar characteristics, or have different characteristics, and then converting this into a quantitative value.
[0220] In this specification, the terms “degree of agreement” and “similarity” used to express homology may be used interchangeably, and specifically, homology may be expressed as degree of agreement.
[0221] For example, if the nucleotides of two nucleic acid sequences are completely identical, their homology is 100%. If there are nucleotides that are not identical between the two nucleic acid sequences, the percentage (%) representing homology decreases. Generally, homology can be a quantitative measure of the degree of identity between two nucleic acid sequences. The degree of homology can be determined by comparing specific positions of each sequence aligned for comparison. If the bases at specific positions of the two sequences being compared are identical, the two nucleic acid sequences are homologous at that position. The degree of homology between two sequences can be calculated as a ratio to the number of homologous positions shared by the two sequences.
[0222] In this specification, the term “align or alignment” refers to a series of techniques for juxtaposing homologous molecular sequences. The alignment and homology values of said sequences may be determined by software known in the art, and various methods and algorithms for alignment are Smith and Waterman, Adv. Appl. Math. 2:482(1981) ; Needleman and Wunsch, J. Mol. Bio. 48:443(1970); Pearson and Lipman, Methods in Mol. Biol. 24: 307-31 (1988); Higgins and Sharp, Gene 73:237-44 (1988); Higgins and Sharp, CABIOS 5:151-3(1989); Corpet et al., Nuc. Acids Res. 16:10881-90 (1988); Huang et al., Comp. Appl. BioSci.8:155-65 (1992) and Pearson et al., Meth. Mol. Biol. It is disclosed in 24:307-31(1994). NCBI Basic Local Alignment Search Tool (BLAST) (Altschul et al., J. Mol. Biol. 215:403-10(1990)) is accessible from sources such as NCBI (National Center for Biological Information) and can be used online in conjunction with sequencing analysis programs such as blastn, blastp, blasm, blastx, tblastn, and tblastx. BLAST can be accessed at http: / / www.ncbi.nlm.nih.gov / BLAST / . Methods for comparing sequence similarity using this program can be found at http: / / www.ncbi.nlm.nih.gov / BLAST / blast_help.html.
[0223] If there is only one taxonomic representative sequence selected in step (c) above, the selected taxonomic representative sequence becomes the group representative sequence.
[0224] If there are two or more taxonomic representative sequences selected in step (c) above, the selected taxonomic representative sequences are grouped according to homology, and a group representative sequence is selected from each group.
[0225] The group representative sequence is a sequence that represents the group, and while the homology for grouping the above-mentioned taxonomic representative sequences is not particularly limited, a homology criterion value may be determined and applied in advance.
[0226] The above homology threshold value may be determined according to the characteristics of the target nucleic acid molecule. For example, it may vary depending on the range of the target nucleic acid molecule. For example, the value of the homology threshold value may vary depending on whether the target nucleic acid molecule is specific to a specific species or specific to a specific subspecies. Alternatively, the above homology threshold value may vary depending on the degree of variation of the target nucleic acid molecule to be detected. Specifically, it may be 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% or more, or 100%; more specifically, the above homology threshold value may be 50%, 60%, 70%, 80%, or 90% or more, or 100%. Alternatively, the above homology threshold value may be determined within the range of 70% to 100%, 80% to 100%, or 90% to 100%.
[0227] The process of grouping the above-mentioned taxonomic representative sequences according to homology and selecting group representative sequences can be carried out by a known program (e.g., UCLUST).
[0228] According to one embodiment of the present invention, the selection of the group representative sequence of step (d) is carried out by a method comprising the following steps:
[0229] (d-1) A step of aligning the selected taxonomic representative sequences to satisfy at least one of the following predetermined alignment criteria;
[0230] (d-2) A step of selecting the highest taxonomic representative sequence among the above-mentioned aligned taxonomic representative sequences; and
[0231] (d-3) grouping taxonomic representative sequences having homology greater than a predetermined value with the above-mentioned top-level taxonomic representative sequence and selecting the above-mentioned top-level taxonomic representative sequence in each group as the group representative sequence, and the alignment criteria include the following:
[0232] (i) Aligned according to the assembly level of the selected taxonomic representative sequences; the assembly levels are ranked in the order of complete genome, chromosome, scaffold, and contig, with the highest ranking.
[0233] (ii) the number of nucleic acid sequence data having the same taxonomic name and / or taxonomic identification symbol as the selected taxonomic representative sequence; the higher the number, the higher the rank, and
[0234] (iii) Sort according to whether a host is listed in the descriptor of the nucleic acid record containing the selected taxonomic representative sequence; the case where the host of interest for the organism of interest is listed in the host is higher in rank than the case where it is not listed, and the case where it is not listed is higher in rank than the case where an organism other than the host of interest is listed in the host, and
[0235] (iv) Sort according to the alphabet of the accession number of the selected taxonomic representative sequence; the earlier the alphabet of the accession number, the higher the rank.
[0236] In an embodiment regarding the selection of taxonomic representative sequences, the description regarding importance and weight among the descriptions of alignment criteria for nucleic acid sequence data having the same taxonomic name and / or taxonomic identifier may be equally applied to alignment criteria related to the selection of group representative sequences.
[0237] The explanation of the homology threshold value used in describing the grouping of taxonomic representative sequences can be applied in the same way to homology greater than a predetermined value in step (d-3) above.
[0238] According to step (d-1) of the present embodiment, the selected taxonomic representative sequence is aligned to satisfy at least one alignment criterion (specifically, alignment criterion (i)), specifically at least two, more specifically at least three, most specifically four, and alignment criteria.
[0239] According to one embodiment of the present invention, the at least two alignment criteria differ in criticality, and the method of the present invention additionally includes the step of aligning taxonomic representative sequences to satisfy the at least two alignment criteria considering the criticality.
[0240] According to step (d-2) of the present embodiment, the highest taxonomic representative sequence among the aligned taxonomic representative sequences is selected.
[0241] According to step (d-3) of the present embodiment, taxonomic representative sequences having homology greater than a predetermined value with the highest taxonomic representative sequence are grouped, and the highest taxonomic representative sequence is selected as the group representative sequence.
[0242] Due to differences in homology among the selected taxonomic representative sequences, the selected taxonomic representative sequences are grouped into multiple groups, and as a result, multiple group representative sequences can be selected.
[0243] When multiple group representative sequences are selected, the homology between the group representative sequences is lower than the homology threshold value used for grouping. If the homology between two group representative sequences is equal to or greater than the homology threshold value, the two group representative sequences are sequences that must belong to the same group.
[0244] When multiple group representative sequences are selected, the multiple group representative sequences may be selected to satisfy the following conditions:
[0245] (i) All group representative sequences shall have homology greater than or equal to the homology threshold value with the taxonomic representative sequences of the group to which they belong; and
[0246] (ii) The homology between the group representative sequences shall have homology less than the homology threshold value of (i) above.
[0247] The explanation of the homology threshold value used in describing the grouping of taxonomic representative sequences can be applied in the same way to homology greater than the homology threshold value of (i) above.
[0248] If multiple taxonomic representative sequences with differences in homology are included to the extent that multiple group representative sequences are selected, the following method may be additionally implemented.
[0249] According to a more specific embodiment of the present invention, the method further comprises the following steps: (d-4) selecting the top taxonomic representative sequence from the remaining taxonomic representative sequences, excluding the group representative sequence and the taxonomic representative sequence included in the group of the group representative sequence from the selected taxonomic representative sequences; and (d-5) performing step (d-3) by replacing the top taxonomic representative sequence of (d-4) with the top taxonomic representative sequence of step (d-3); and, if there are additional taxonomic representative sequences to be grouped, repeat steps (d-4) and (d-5).
[0250] Step (e): Provides nucleic acid sequence data sets for oligonucleotide design ( 150 )
[0251] Finally, the method of the present invention (e) collects nucleic acid sequence data having homology greater than a predetermined value with the group representative sequence and provides it as a nucleic acid sequence data set for designing the oligonucleotide.
[0252] According to the present invention, nucleic acid sequence data is collected from the name and / or synonyms of a target nucleic acid molecule of an organism of interest, a taxonomic representative sequence and a group representative sequence are selected therefrom, and then nucleic acid sequence data having homology greater than a predetermined value with the group representative sequence is collected and provided as a nucleic acid sequence data set for designing the oligonucleotide.
[0253] If there are two or more of the above group representative sequences, nucleic acid sequence data having homology of a predetermined value or more for each group representative sequence is collected and provided as a nucleic acid sequence data set for designing the above oligonucleotides.
[0254] The above-mentioned homology value is not specifically limited, but, for example, may be 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or more, or 100% homology. Specifically, it may be 40%, 50%, 60%, 70%, 80%, 90% or more, or 100% homology, and more specifically, it may be 40%, 50%, 60%, 70%, or 80% or more. Alternatively, it may be selected within a range of 10% to 100%, more specifically within a range of 40% to 100%, and even more specifically within a range of 40%, 50%, 60%, 70%, or 80%.
[0255] According to one embodiment of the present invention, nucleic acid sequence data having homology greater than a predetermined value is collected from a second database.
[0256] The description of the second database in step (b) above may be applied equally to this step, and the common content between them is omitted to avoid excessive complexity of this specification due to repetition.
[0257] The above collection can be performed using software known in the industry (e.g., BLAST).
[0258] The above provision may be implemented by various data provision methods known in the art. For example, it may be provided by exposing the data content to the user in a state where it can be directly perceived through an output device or a display device, or by storing the data on a data storage medium intended by the user through a recording device, or by transmitting the data to an intended device through a network device capable of wired or wireless data transmission.
[0259] According to one embodiment of the present invention, the nucleic acid sequence data set for designing the provided oligonucleotide is a list of nucleic acid sequence data sets including information regarding the nucleic acid sequence data set and the nucleic acid sequence data set, and / or an alignment result in which the nucleic acid sequence data set is aligned.
[0260] The information regarding the above nucleic acid sequence data represents information including all information described in the nucleic acid record containing the above nucleic acid sequence data, and includes, for example, the access number of the above nucleic acid sequence data, the group number containing the above nucleic acid sequence data, the location information of the gene in the above nucleic acid sequence data, the name of the organism (or taxonomic name), the taxonomic ID, the name of the gene, the name of the protein, homology information, etc.
[0261] The alignment result of the above nucleic acid sequence data set is an alignment result aligned using various alignment programs known in the art, and the alignment result is provided in the form of a file provided by the alignment program.
[0262] The descriptions of various methods and algorithms for alignment in step (c) above are equally applicable to this step, and common details among them are omitted to avoid excessive complexity in this specification due to repetition.
[0263] According to one embodiment of the present invention, the method further comprises the following step after step (e):
[0264] (e-1) A step of sorting the design nucleic acid sequence data set provided above according to biosample identifiers to select nucleic acid sequence data having the same biosample identifier;
[0265] (e-2) A step of aligning the selected nucleic acid sequence data to satisfy at least one of the following alignment criteria;
[0266] (e-3) A step of selecting the nucleic acid sequence of the highest nucleic acid sequence data among the above-mentioned aligned nucleic acid sequence data; and
[0267] (e-4) A step of removing nucleic acid sequence data other than the top-most nucleic acid sequence data from the design nucleic acid sequence data set, and the alignment criteria include the following:
[0268] (i) The nucleic acid sequences included in the design nucleic acid sequence data set provided above are aligned according to the assembly level; the assembly levels are ranked in the order of complete genome, chromosome, scaffold, and contig (i.e., the complete genome has the highest rank),
[0269] (ii) Sort the nucleic acid sequences included in the design nucleic acid sequence data set provided above based on whether they are included in the RefSeq (Reference Sequence) database; if the nucleic acid sequence data is included in the RefSeq database, the ranking is higher than if it is not included.
[0270] In this embodiment, the description of steps (b-3) to (b-6) performed between steps (b) and (c) can be applied identically, and common details between the two embodiments are omitted to avoid excessive complexity in this specification due to repetition.
[0271] The removal of duplicate sequences carried out by steps (b-3) to (b-6) above is carried out on nucleic acid sequence data collected as names of target nucleic acid molecules for organisms of interest and / or synonyms thereof, but the removal of duplicate sequences according to the present embodiment is carried out on the nucleic acid sequence data set for designing oligonucleotides provided in step (e).
[0272] In addition, the present embodiment can remove duplicate sequences by using the target nucleic acid sequence data set provided in step (f) or the non-target nucleic acid sequence data set provided in step (g) described below as targets for removing duplicate sequences.
[0273] By removing duplicate sequences according to this embodiment, aspects of coverage and costs associated with the introduction of degenerate bases can be improved in designing oligonucleotides.
[0274] A nucleic acid sequence data set for the design of an oligonucleotide provided by the method of the present invention comprises nucleic acid sequence data relating to the received organism of interest and / or nucleic acid sequence data not relating to the received organism of interest.
[0275] In this specification, among the nucleic acid sequence data sets for designing oligonucleotides, the nucleic acid sequence data set relating to an organism of interest represents a target nucleic acid sequence data set relating to a target nucleic acid molecule, and among the nucleic acid sequence data sets for designing oligonucleotides, the nucleic acid sequence data set relating not to an organism of interest represents a non-target nucleic acid sequence data set relating to a non-target nucleic acid molecule.
[0276] Oligonucleotides used to amplify or detect target nucleic acid molecules of an organism of interest must satisfy the following two design requirements: First, they must be able to detect multiple target nucleic acid sequences having sequence similarity to the target nucleic acid molecule of the organism of interest with the highest possible target coverage. Second, they must not detect nucleic acid molecules of organisms other than the organism of interest.
[0277] In order to design an oligonucleotide satisfying these two requirements, among the nucleic acid sequence data sets provided, for the first requirement, a target nucleic acid sequence data set for a target nucleic acid molecule is provided, and for the second requirement, a non-target nucleic acid sequence data set for a non-target nucleic acid molecule is provided.
[0278] Step (f): Provides target nucleic acid sequence data sets for target nucleic acid molecules
[0279] According to one embodiment of the present invention, the method further comprises the step of (f) providing the nucleic acid sequence data relating to the received organism of interest among the nucleic acid sequence data sets provided in step (e) as a target nucleic acid sequence data set relating to a target nucleic acid molecule.
[0280] Collecting target nucleic acid sequences using the name of the target nucleic acid molecule of the organism of interest and / or its synonyms is impossible if the name of the nucleic acid sequence is changed after registration, or if information regarding the nucleic acid sequence is incorrectly entered or omitted due to an error by the sequence registrant. A method of providing a target nucleic acid sequence data set regarding the received organism of interest from among nucleic acid sequence data having homology of a predetermined value or greater with a group representative sequence can resolve the problem of target nucleic acid sequence data being omitted for the reasons mentioned above.
[0281] The nucleic acid sequence data regarding the received organism of interest refers to nucleic acid sequence data in which the name or synonym thereof of the organism of interest received in step (a), or the name or synonym thereof of an organism belonging to a taxonomic subcategory of the received organism of interest, is listed as the organism in the nucleic acid sequence data. However, the nucleic acid sequence data regarding the received organism of interest does not refer to nucleic acid sequence data in which the name or synonym thereof of an organism belonging to a taxonomic supercategory of the received organism of interest is listed as the organism in the nucleic acid sequence data.
[0282] According to the present embodiment, after collecting nucleic acid sequence data having homology greater than a predetermined value with the group representative sequence from a nucleotide database, information regarding the organism of the collected sequence is compared with information regarding the organism of interest of the target nucleic acid molecule to provide a target nucleic acid sequence data set for the target nucleic acid molecule.
[0283] The information regarding the above organism may be the scientific name or taxonomic name of the organism listed in 'organism' as the title or descriptor of the nucleic acid record for each nucleic acid sequence, or the taxonomic ID assigned to the scientific name or taxonomic name of the organism.
[0284] The above provision may provide the collected nucleic acid sequence data as nucleic acid sequence data included in the target nucleic acid sequence data set for the target nucleic acid molecule if the information regarding the organism of interest of the target nucleic acid molecule is identical to the information regarding the organism of interest of the target nucleic acid molecule, or if it is identical to the information regarding the sub-organism of the organism of interest of the target nucleic acid molecule.
[0285] According to the present embodiment, sequences that are not included in name-based sequence collection can be collected because some synonyms are omitted during the synonym collection process or the name of the target nucleic acid molecule is incorrectly written when the initial sequence is registered.
[0286] In this specification, “target nucleic acid sequence” refers to a sequence associated with a target nucleic acid molecule, which is a nucleic acid molecule to be finally detected. The target nucleic acid sequence may include the entire or a part thereof of the nucleic acid sequence corresponding to the target nucleic acid molecule.
[0287] The nucleic acid sequences for a common specific gene possessed by a specific group of organisms may be identical or different for each individual. Therefore, when a target nucleic acid molecule represents a common specific gene possessed by a specific group of organisms, the nucleic acid sequence corresponding to the target nucleic acid molecule cannot be determined as a single sequence of nucleotides. In other words, for a single target nucleic acid molecule, there may exist multiple target nucleic acid sequence data with different sequences of nucleotides.
[0288] A target nucleic acid sequence data set refers to a collection of target nucleic acid sequence data. In other words, a target nucleic acid sequence data set refers to a collection of information regarding the sequence of nucleotides of a target nucleic acid molecule. As described above, for a single target nucleic acid molecule, there may exist various target nucleic acid sequence data with different sequences of nucleotides.
[0289] Accordingly, according to one embodiment of the present invention, the target nucleic acid sequence data set for the target nucleic acid molecule may be a data set comprising a plurality of target nucleic acid sequence data.
[0290] According to one embodiment of the present invention, the target nucleic acid sequence data set may include nucleic acid sequence data corresponding to part or all of the target nucleic acid molecule, or variant nucleic acid sequence data for the target nucleic acid molecule.
[0291] Since target nucleic acid sequence data for a target nucleic acid molecule refers to nucleic acid sequence data related to the nucleic acid sequence of a target nucleic acid molecule of an organism of interest, the target nucleic acid sequence data includes both nucleic acid sequence data consisting of all or part of the nucleic acid sequence corresponding to the target nucleic acid molecule and nucleic acid sequence data including all or part of the nucleic acid sequence corresponding to the target nucleic acid molecule of an organism of interest.
[0292] A variant nucleic acid sequence for a target nucleic acid molecule refers to a nucleic acid sequence comprising a nucleotide sequence in which one or more nucleotides are substituted, deleted, and / or added compared to the target nucleic acid sequence of the target nucleic acid molecule.
[0293] According to one embodiment of the present invention, a target nucleic acid sequence dataset for the target nucleic acid molecule may be provided including information aligned with a group representative sequence. More specifically, the target nucleic acid sequence dataset for the target nucleic acid molecule may be provided in the form of an alignment file in which the target nucleic acid sequences of the target nucleic acid sequence dataset are aligned with the group representative sequence. Through this aligned information, the target nucleic acid sequence dataset for the target nucleic acid molecule can be used more effectively for the design of an oligonucleotide.
[0294] According to a more specific embodiment of the present invention, the target nucleic acid sequence data set provided in step (f) has homology of at least a predetermined value with respect to the target nucleic acid sequence data set with respect to at least one representative sequence among the group representative sequence and the taxonomic representative sequence.
[0295] According to the present embodiment, in step (e), nucleic acid sequence data having homology greater than or equal to a predetermined value with respect to a group representative sequence is collected, and in step (f), nucleic acid sequence data relating to an organism of interest received in step (a) among the collected nucleic acid sequence data is provided, and then, with respect to the nucleic acid sequence data, nucleic acid sequence data having homology greater than or equal to a predetermined value with respect to at least one representative sequence among the group representative sequence and the taxonomic representative sequence can be provided as a target nucleic acid sequence data set. Accordingly, the present embodiment can also be expressed as the time-series process described above.
[0296] A predetermined value for homology in the present embodiment is greater than a predetermined value for homology in step (e), and specifically, a predetermined value for homology in the present embodiment may be 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, or 60% greater than a predetermined value for homology in step (e).
[0297] In this embodiment, unlike in step (e), for the nucleic acid sequence data relating to the organism of interest among the nucleic acid sequence data sets having homology with the group representative sequence, a nucleic acid sequence data set having homology with the group representative sequence and / or the taxonomic representative sequence is additionally provided as a target nucleic acid sequence data set.
[0298] In this embodiment, the reason for considering not only the group representative sequence but also the taxonomic representative sequence is to collect the target nucleic acid sequence data that the oligonucleotide must detect without omission when designing the oligonucleotide used to detect the target nucleic acid molecule.
[0299] The predetermined value of homology in the present embodiment may be 70%, 80%, 90%, 95%, 96%, 97%, 98%, or 99% or more, or 100% homology. The predetermined value may be a value within a certain range, but is not limited thereto; for example, it may be 70% to 100%, 80% to 100%, 90% to 100%, or 95% to 100%.
[0300] Step (g): Provides non-target nucleic acid sequence datasets for non-target nucleic acid molecules
[0301] According to one embodiment of the present invention, the method further comprises the step of (g) providing, among the nucleic acid sequence data sets provided in step (e), nucleic acid sequence data that are not related to the received organism of interest as a non-target nucleic acid sequence data set for non-target nucleic acid molecules.
[0302] According to the present embodiment, among the nucleic acid sequence data sets having homology greater than a predetermined value with the group representative sequence collected in step (e), the nucleic acid sequence data that is not related to the received organism of interest is referred to as a non-target nucleic acid sequence data set for non-target nucleic acid molecules.
[0303] This embodiment is implemented to meet the second requirement among the aforementioned oligonucleotide design requirements, which is that nucleic acid molecules of an organism other than the organism of interest must not be detected. Specifically, another issue with the design of oligonucleotides for detecting specific target nucleic acid molecules is that information regarding nucleic acid molecules that may cause false positive errors must be identified to ensure that such nucleic acid molecules are not detected.
[0304] To solve these problems, the present embodiment provides a non-target nucleic acid sequence dataset for non-target nucleic acid molecules that may cause false positive errors, and can be used to design oligonucleotides so that the non-target nucleic acid sequence dataset is not detected.
[0305] The nucleic acid sequence data not relating to the received organism of interest refers to nucleic acid sequence data in which the name or synonym of the organism of interest received in step (a), or the name or synonym of an organism belonging to a taxonomic subcategory of the received organism of interest, is not listed as an organism in the nucleic acid sequence data. Accordingly, the nucleic acid sequence data not relating to the received organism of interest refers to nucleic acid sequence data in which the name or synonym of an organism belonging to a taxonomic supercategory of the received organism of interest, or the name or synonym of an organism different from the organism of interest, is listed as an organism in the nucleic acid sequence data.
[0306] According to the present embodiment, after collecting nucleic acid sequence data having homology greater than a predetermined value with the group representative sequence from a nucleotide database, information regarding the organism of the collected sequence is compared with information regarding the organism of interest of the target nucleic acid molecule to provide a non-target nucleic acid sequence data set for the non-target nucleic acid molecule.
[0307] Specifically, nucleic acid sequence data having homology greater than a predetermined value and information regarding the organism of interest thereof are collected from a second database, and among the collected nucleic acid sequence data, nucleic acid sequence data for which information regarding the organism does not belong to the organism of interest or its sub-organism received in step (a) may be provided as a non-target nucleic acid sequence for a non-target nucleic acid molecule.
[0308] The information regarding the above organism may be the scientific name or taxonomic name of the organism listed in 'organism' as the title or descriptor of the nucleic acid record for each nucleic acid sequence, or the taxonomic ID assigned to the scientific name or taxonomic name of the organism.
[0309] The above provision may provide the information regarding the organism in the collected nucleic acid sequence data as nucleic acid sequence data included in a non-target nucleic acid sequence data set for a non-target nucleic acid molecule if the information regarding the organism of interest of the target nucleic acid molecule is different from the information regarding the organism of interest of the target nucleic acid molecule, or if it is identical to the information regarding the parent organism of the organism of interest of the target nucleic acid molecule.
[0310] In this specification, the term “non-target nucleic acid molecule” refers to a nucleic acid molecule that is the opposite of a target nucleic acid molecule and must not be detected during the detection process of the target nucleic acid molecule, regardless of homology with the sequence of the target nucleic acid molecule. In this specification, the term “non-target nucleic acid sequence” refers to the nucleic acid sequence of a non-target nucleic acid molecule.
[0311] According to one embodiment of the present invention, a non-target nucleic acid sequence data set for the non-target nucleic acid molecule may be provided including information aligned with a group representative sequence. More specifically, the non-target nucleic acid sequence data set for the non-target nucleic acid molecule may be provided in the form of an alignment file in which the non-target nucleic acid sequences of the non-target nucleic acid sequence data set are aligned with the group representative sequence.
[0312] The non-target nucleic acid sequence dataset for non-target nucleic acid molecules provided by this method contains sequences similar to the target nucleic acid molecule but includes nucleic acid sequences that are not the target nucleic acid sequence. Therefore, when designing or providing an oligonucleotide for detecting a target nucleic acid molecule, if the design is made so as not to hybridize with the nucleic acid sequences included in the non-target nucleic acid sequence dataset for the non-target nucleic acid molecule, it is possible to provide an oligonucleotide for detecting a target nucleic acid molecule with high specificity and no risk of false positives.
[0313] According to one embodiment of the present invention, the non-target nucleic acid sequence data set provided in step (g) satisfies at least one of the following homology criteria:
[0314] (i) The above non-target nucleic acid sequence data set shall have homology greater than a predetermined value with respect to a portion of the sequence region of at least one of the group representative sequence and the taxonomic representative sequence;
[0315] (ii) The above non-target nucleic acid sequence data set shall have homology greater than a predetermined value with respect to at least one representative sequence among the group representative sequence and the taxonomic representative sequence; and
[0316] (iii) A non-target nucleic acid sequence data set having the homology criterion of (i) above will have the homology criterion of (ii) above.
[0317] According to the present embodiment, in step (e), nucleic acid sequence data having homology greater than a predetermined value with respect to a group representative sequence is collected, and in step (g), nucleic acid sequence data among the collected nucleic acid sequence data that does not relate to the organism of interest received in step (a) is provided, and then nucleic acid sequence data having homology greater than a predetermined value with respect to at least one representative sequence among the group representative sequence and the taxonomic representative sequence among the nucleic acid sequence data can be provided as a non-target nucleic acid sequence data set. Accordingly, the present embodiment can also be expressed as the time-series process described above.
[0318] According to the homology criterion (i) in the present embodiment, the non-target nucleic acid sequence data set is required to have homology of at least a predetermined value with respect to a portion of the sequence region of at least one representative sequence among the group representative sequence and the taxonomic representative sequence.
[0319] A partial sequence region of at least one of the group representative sequences and taxonomic representative sequences refers to a non-target nucleic acid sequence included in the non-target nucleic acid sequence data set aligned for homology comparison with the at least one representative sequence, and a region having a certain nucleotide length from one end of the at least one representative sequence, and specifically, the nucleotide length representing the partial sequence region is 10bp, 20bp, 30bp, 40bp, 50bp, 60bp, or 70bp, but is not limited thereto.
[0320] A predetermined value of homology in the above sequence region is desirable to be greater than the predetermined value of homology in the above sequence region, as a design requirement to design the oligonucleotide so as not to detect non-target nucleic acid sequences with high homology in the above sequence region. Specifically, it is 80%, 90%, 95%, 98%, or 99% or more, or 100%.
[0321] According to the homology criterion (ii) in the present embodiment, the non-target nucleic acid sequence data set is required to have homology greater than a predetermined value with respect to at least one representative sequence among the group representative sequence and the taxonomic representative sequence.
[0322] According to the above homology criterion (ii), the non-target nucleic acid sequence data set is required to have homology greater than a predetermined value with respect to at least one representative sequence.
[0323] Since the homology criterion (ii) is compared with at least one representative sequence, the predetermined value of the homology may be lower than the homology criterion (i). Specifically, the predetermined value of the homology criterion (ii) is 60%, 70%, 80%, 90%, or 95% or more, or 100%.
[0324] According to homology criterion (iii) in the present embodiment, it is required to have both of the homology criteria of (i) and (ii).
[0325] In this embodiment, unlike in step (e), for nucleic acid sequence data that is not related to the organism of interest among nucleic acid sequence data sets having homology with the group representative sequence, a nucleic acid sequence data set having homology with the group representative sequence and / or the taxonomic representative sequence is additionally provided as a non-target nucleic acid sequence data set.
[0326] In this embodiment, the reason for considering not only the group representative sequence but also the taxonomic representative sequence is to collect non-target nucleic acid sequence data that the oligonucleotide must not detect without omission when designing the oligonucleotide used to detect the target nucleic acid molecule.
[0327] The reason for requiring a homology criterion for non-target nucleic acid sequences in this embodiment is that, as described above, the problem of false positive errors caused by non-target nucleic acid molecules similar to the target nucleic acid molecule becomes more problematic when a portion of the non-target nucleic acid molecule's sequence exhibits very high similarity to the sequence of the target nucleic acid molecule.
[0328] Generally, non-target nucleic acid sequences in which only specific regions exhibit very high homology with the target nucleic acid sequence while other regions have low homology, resulting in low overall sequence homology with the target nucleic acid molecule, may not be considered during the design process of oligonucleotides for detecting the target nucleic acid molecule. Accordingly, to solve this problem, the method of the present invention separately selects non-target sequences that have somewhat low overall homology with the target nucleic acid molecule but high homology in certain regions, and provides information regarding such sequences.
[0329] Steps (h) and (j): Classified as target nucleic acid sequence data excluded from design
[0330] According to one embodiment of the present invention, the method further comprises the following steps:
[0331] (h) a step of collecting nucleic acid sequence data having homology greater than or equal to a predetermined value with the above group representative sequence and providing a nucleic acid sequence data set; and
[0332] (j) a step of classifying the target nucleic acid sequence data of the group representative sequence and the target nucleic acid sequence data belonging to the same group as the group representative sequence in the target nucleic acid sequence data set of step (f) as design-excluded target nucleic acid sequence data, provided that the taxonomic name and / or taxonomic identification symbol of the organism for the group representative sequence satisfies one of the following predetermined criteria; and, the predetermined criteria include the following:
[0333] (i) If the nucleic acid sequence data set provided in step (h) above does not contain nucleic acid sequence data corresponding to the taxonomic name and / or taxonomic identification symbol of the organism for the group representative sequence, and only nucleic acid sequence data corresponding to the taxonomic name and / or taxonomic identification symbol of the organism and other organisms for the group representative sequence exists;
[0334] (ii) If the homology of the nucleic acid sequence data corresponding to the taxonomic name and / or taxonomic identification symbol of the organism for the group representative sequence in the nucleic acid sequence data set provided in step (h) is lower compared to the homology of the nucleic acid sequence data corresponding to the taxonomic name and / or taxonomic identification symbol of the organism for the group representative sequence and of other organisms;
[0335] (iii) In the nucleic acid sequence data set provided in step (h), there is no nucleic acid sequence data corresponding to the taxonomic name and / or taxonomic identification symbol of an organism for the group representative sequence, and the ratio of nucleic acid sequence data corresponding to an organism corresponding to a taxonomic name and / or taxonomic identification symbol higher than that of the organism for the group representative sequence, or an organism corresponding to a taxonomic name and / or taxonomic identification symbol lower than that, is less than a predetermined value for the nucleic acid sequence data set provided in step (h); and
[0336] (iv) where the nucleic acid sequence data set provided in step (h) above is a nucleic acid sequence data set of an organism for the group representative sequence, but the name of the target nucleic acid molecule is not listed in the descriptors of the nucleic acid records containing the nucleic acid sequence data set, or is different from at least one of the name of the target nucleic acid molecule and the collected synonyms.
[0337] According to the present invention, a target nucleic acid sequence data set for a target nucleic acid molecule is a nucleic acid sequence data set that must be considered such that the oligonucleotide used to amplify or detect the target nucleic acid molecule of an organism of interest hybridizes to the target nucleic acid sequence data set.
[0338] However, when designing to hybridize to all of the above target nucleic acid sequence data sets, the following problems may arise. If there are registration errors in the nucleotide database among the collected target nucleic acid sequence data (cases where an organism is registered in the database as the same organism as the organism of interest, but the registered nucleic acid sequence is a nucleic acid sequence of a different organism than the organism of interest), and an attempt is made to design an oligonucleotide to hybridize to all of these registration errors, not only is the target coverage of the oligonucleotide lowered, but there are also problems such as introducing degenerate bases into the oligonucleotide or increasing the number of combinations in order to increase the target coverage.
[0339] Accordingly, as in the present embodiment, it is necessary to classify target nucleic acid sequence data that is excluded from the design of oligonucleotides among the target nucleic acid sequence data sets for target nucleic acid molecules by checking for errors in the registration of group representative sequences.
[0340] According to one embodiment of the present invention, nucleic acid sequence data having homology of step (h) is collected from a third database.
[0341] The third database used to collect homologous nucleic acid sequence data in step (h) of the present embodiment may be the same as the second database described above or may be a nucleotide database that includes a part of the nucleotide database of the second database.
[0342] In the case where the third database includes some nucleotide databases of the second database, the third database may be a nucleotide database including NCBI’s GenBank (including SNP and non-WGS databases), RefSeq, DDBJ, and EMBL databases, or a nucleotide database built by downloading the nucleotide database.
[0343] Although step (h) of the present embodiment is described as collecting nucleic acid sequence data having homology greater than a predetermined value with respect to the group representative sequence (specifically, from the third database), the nucleic acid sequence data collected and provided in step (e) (specifically, from the second database) may be used as the nucleic acid sequence data set provided in step (h). Specifically, if the third database and the second database are the same, for the nucleic acid sequence data among the nucleic acid sequence data collected in step (e) that satisfies the homology criteria of step (h), it is checked whether the taxonomic name and / or taxonomic identification symbol of the organism for the group representative sequence satisfies one of the predetermined criteria.
[0344] Alternatively, if the third database includes a portion of the nucleotide database of the second database, the nucleic acid sequence data corresponding to the third database among the nucleic acid sequence data collected in step (e) is collected, and for the nucleic acid sequence data among the collected nucleic acid sequence data that satisfies the homology criteria of step (h), it is checked whether the taxonomic name and / or taxonomic identification symbol of the organism for the group representative sequence satisfies one of the predetermined criteria.
[0345] In step (j) of the present embodiment, the results of collecting nucleic acid sequence data homologous thereto are compared with a selected group representative sequence, and if the taxonomic name and / or taxonomic identification symbol of the organism for the group representative sequence satisfies one of the predetermined criteria, the target nucleic acid sequence data of the group representative sequence and the target nucleic acid sequence data belonging to the same group as the group representative sequence are classified as design-excluded target nucleic acid sequence data in the target nucleic acid sequence data set of step (f).
[0346] If a target nucleic acid sequence is classified as a design-excluded target nucleic acid sequence among the target nucleic acid sequence data, the oligonucleotide is not designed by referencing the design-excluded target nucleic acid sequence data. Therefore, the design-excluded target nucleic acid sequence data is excluded from the target nucleic acid sequence data, and the oligonucleotide is designed based on it. In other words, the design-excluded target nucleic acid sequence becomes a nucleic acid sequence that may or may not be detected as an oligonucleotide.
[0347] The predetermined value of homology considered in step (h) above is greater than the homology considered in step (e), but may be equal to the homology criterion of the target nucleic acid sequence data set in step (f).
[0348] Specifically, the predetermined value of homology in step (h) above may be 70%, 80%, 90%, 95%, 96%, 97%, 98%, or 99% or more, or 100% homology. The predetermined value may be a value within a certain range, but is not limited thereto; for example, it may be 70% to 100%, 80% to 100%, 90% to 100%, or 95% to 100%.
[0349] The ratio of nucleic acid sequence data of the criterion (iii) of step (j) above may be less than 2%, 5%, 8%, 10%, 15%, 20%, 25%, 30%, 35%, or 40%.
[0350] Among the target nucleic acid sequence data remaining after excluding design-excluded nucleic acid sequence data, multiple target nucleic acid sequence data having the same access number may exist within one or multiple groups. In this case, nucleic acid sequence data among the multiple target nucleic acid sequence data where part or all of the nucleic acid sequence data overlaps may be merged and provided.
[0351] According to one embodiment of the present invention, the method may additionally include, after step (j), a step of (j-1) providing target nucleic acid sequence data by merging target nucleic acid sequence data in which part or all of the nucleic acid sequence data overlaps among the target nucleic acid sequence data remaining after excluding the design-excluded nucleic acid sequence data among the target nucleic acid sequence data provided in step (f), when there are multiple target nucleic acid sequence data having the same access number within a plurality of groups.
[0352] For example, target nucleic acid sequence data with the same access number may exist for multiple groups due to differences in regions homologous to multiple group representative sequences, and consequently, the target nucleic acid sequence data set may contain duplicate sequences. In such cases, it is not appropriate to use each of the multiple target nucleic acid sequence data with the same access number as the target nucleic acid sequence for the target nucleic acid molecule; instead, it is preferable to treat a single nucleic acid sequence formed by merging them as the target nucleic acid sequence for the target nucleic acid molecule.
[0353] According to one embodiment of the present invention, the method further comprises, after step (j-1), step (j-2) of comparing the homology of the merged target nucleic acid sequence data with the group representative sequence within the plurality of groups and including the group representative sequence with high homology in the group to which it belongs.
[0354] The present embodiment is a process of merging multiple target nucleic acid sequence data included in multiple groups, and then determining the group of a single merged target nucleic acid sequence data.
[0355] In this way, by merging partially or wholly overlapping target nucleic acid sequence data into a single target nucleic acid sequence data, target nucleic acid sequence data for parts of the target nucleic acid molecule are recognized as target nucleic acid sequence data for independent target nucleic acid molecules, thereby reducing the probability of statistical error occurring in analysis based on the target nucleic acid sequence data set.
[0356] Steps (k) and (l): Classified as non-target nucleic acid sequence data, excluded from design
[0357] According to one embodiment of the present invention, the method further comprises the following steps:
[0358] (k) a step of collecting nucleic acid sequence data having homology greater than or equal to a predetermined value with respect to the non-target nucleic acid sequence of an organism for the above non-target nucleic acid sequence data set and providing a nucleic acid sequence data set; and
[0359] (l) a step of classifying the non-target nucleic acid sequence data of the organism in the non-target nucleic acid sequence data set of step (k) as design-excluded non-target nucleic acid sequence data, provided that the taxonomic name and / or taxonomic identification symbol of the organism for the non-target nucleic acid sequence satisfies one of the following predetermined criteria; and, the predetermined criteria include the following:
[0360] (i) If the nucleic acid sequence data set provided in step (k) above does not contain nucleic acid sequence data corresponding to the taxonomic name and / or taxonomic identification symbol of the organism for the non-target nucleic acid sequence, and only nucleic acid sequence data corresponding to the taxonomic name and / or taxonomic identification symbol of the organism and other organisms exists;
[0361] (ii) if the homology of the non-target nucleic acid sequence data corresponding to the taxonomic name and / or taxonomic identification symbol of the organism for the non-target nucleic acid sequence in the nucleic acid sequence data set provided in step (k) is lower compared to the homology of the nucleic acid sequence data corresponding to the taxonomic name and / or taxonomic identification symbol of the organism and other organisms; and
[0362] (iii) In the nucleic acid sequence data set provided in step (k), there is no non-target nucleic acid sequence data corresponding to the taxonomic name and / or taxonomic identification symbol of the organism for the non-target nucleic acid sequence, and the ratio of nucleic acid sequence data corresponding to an organism corresponding to a taxonomic name and / or taxonomic identification symbol higher than the organism or an organism corresponding to a taxonomic name and / or taxonomic identification symbol lower than the same is less than a predetermined value for the nucleic acid sequence data set provided in step (k).
[0363] According to the present invention, a non-target nucleic acid sequence data set for a non-target nucleic acid molecule is a nucleic acid sequence data set that must be considered so that the oligonucleotide used to amplify or detect a target nucleic acid molecule of an organism of interest does not necessarily hybridize to the non-target nucleic acid sequence data set.
[0364] However, when designed not to hybridize to all of the above non-target nucleic acid sequence data sets, the following problems may arise. Among the collected non-target nucleic acid sequence data, there may be errors in the registration of the nucleotide database (e.g., a case where an organism A is registered in the database for a non-target nucleic acid sequence, but the registered nucleic acid sequence is a nucleic acid sequence of an organism different from said organism A), or there may be cases where it is difficult to design the oligonucleotide so as not to detect said collected non-target nucleic acid sequences due to high homology with the group representative sequence.
[0365] Accordingly, as in the present embodiment, it is necessary to identify errors in the registration of organisms in the collected non-target nucleic acid sequence data, and to determine a non-target nucleic acid sequence data set that is not considered in the design of oligonucleotides, i.e., a non-target nucleic acid sequence data set excluded from the design.
[0366] According to one embodiment of the present invention, nucleic acid sequence data having homology of step (k) is collected from a third database.
[0367] The third database used to collect homologous nucleic acid sequence data in step (k) of the present embodiment may be the same as the second database described above, or a nucleotide database including a part of the nucleotide database of the second database.
[0368] In the case where the third database includes some nucleotide databases of the second database, the third database may be a nucleotide database including NCBI’s GenBank (including SNP and non-WGS databases), RefSeq, DDBJ, and EMBL databases, or a nucleotide database built by downloading the nucleotide database.
[0369] According to one embodiment of the present invention, the non-target nucleic acid sequence data set of step (k) is a non-target nucleic acid sequence data set having homology of a group representative sequence of the non-target nucleic acid sequence data set provided in step (g) at a value greater than a predetermined value, and the predetermined value of homology may be the same as the predetermined value of homology of the target nucleic acid sequence data set provided in step (f). Specifically, the homology greater than or equal to the predetermined value may be 70%, 80%, 90%, 95%, 96%, 97%, 98%, or 99% or greater, or 100% homology. The predetermined value may be a value within a certain range, but is not limited thereto; for example, it may be 70% to 100%, 80% to 100%, 90% to 100%, or 95% to 100%.
[0370] The predetermined value of the homology of the nucleic acid sequence data collected in step (k) of the present embodiment may be 70%, 80%, 90%, 95%, 96%, 97%, 98%, or 99% or more, or 100% homology. The predetermined value may be a value within a certain range, but is not limited thereto; for example, it may be 70% to 100%, 80% to 100%, 90% to 100%, or 95% to 100%.
[0371] In step (l) of the present embodiment, the results of collecting the non-target nucleic acid sequence of an organism for the non-target nucleic acid sequence data set and the nucleic acid sequence data homologous thereto are compared, and if the taxonomic name and / or taxonomic identification symbol of the organism for the non-target nucleic acid sequence satisfies one of the predetermined criteria, the non-target nucleic acid sequence data of the organism in the non-target nucleic acid sequence data set of step (g) is classified as design-excluded non-target nucleic acid sequence data.
[0372] If non-target nucleic acid sequence data is classified as non-target nucleic acid sequence data excluded from the design, the oligonucleotide is not designed by referencing the non-target nucleic acid sequence data excluded from the design. Therefore, the non-target nucleic acid sequence data excluded from the design is excluded from the non-target nucleic acid sequence data, and the oligonucleotide is designed so as not to hybridize with the remaining non-target nucleic acid sequence data. In other words, the non-target nucleic acid sequence excluded from the design becomes a nucleic acid sequence that may or may not be detected by the oligonucleotide.
[0373] The ratio of nucleic acid sequence data of the criterion (iii) of step (l) above may be less than 2%, 5%, 8%, 10%, 15%, 20%, 25%, 30%, 35%, or 40%.
[0374] Recording media, devices, and programs
[0375] According to another aspect of the present invention, the present invention comprises a computer-readable recording medium including instructions for implementing a processor for executing a method for providing a nucleic acid sequence data set for designing an oligonucleotide used to detect a target nucleic acid molecule of an organism of interest, wherein the method comprises the following steps: (a) receiving the name of the target nucleic acid molecule and the name of the organism of interest, and retrieving synonyms for the target nucleic acid molecule of the organism of interest; (b) collecting nucleic acid sequence data contained in nucleic acid records; each of the nucleic acid records relates to the organism of interest and includes a descriptor in which at least one of the name of the target nucleic acid molecule and the collected synonyms is described; and (c) sorting the collected nucleic acid sequence data according to a taxonomic name and / or taxonomic ID and selecting a taxonomic representative sequence from among the nucleic acid sequence data having the same taxonomic name and / or taxonomic ID. (d) grouping the selected taxonomic representative sequences according to homology and selecting a group representative sequence from each group; and (e) collecting nucleic acid sequence data having homology greater than a predetermined value with the group representative sequence and providing it as a nucleic acid sequence data set for designing the oligonucleotide.
[0376] According to another aspect of the present invention, the present invention provides a computer program stored on a computer-readable recording medium, which implements a processor for executing a method for providing a nucleic acid sequence data set for designing an oligonucleotide used to detect a target nucleic acid molecule of an organism of interest, said method comprising the following steps: (a) receiving the name of the target nucleic acid molecule and the name of the organism of interest, and retrieving synonyms for said target nucleic acid molecule of said organism of interest; (b) collecting nucleic acid sequence data contained in nucleic acid records; Each of the above nucleic acid records relates to the organism of interest and includes a descriptor in which the name of the target nucleic acid molecule and at least one of the collected synonyms are listed; (c) sorting the collected nucleic acid sequence data according to a taxonomic name and / or a taxonomic ID and selecting a taxonomic representative sequence from among the nucleic acid sequence data having the same taxonomic name and / or taxonomic ID; (d) grouping the selected taxonomic representative sequences according to homology and selecting a group representative sequence from each group; and (e) collecting nucleic acid sequence data having homology of a predetermined value or more with the group representative sequence and providing it as a nucleic acid sequence data set for designing the oligonucleotide.
[0377] According to another aspect of the present invention, the present invention provides an apparatus for providing a nucleic acid sequence data set for designing oligonucleotides used to detect target nucleic acid molecules of an organism of interest, comprising (a) a computer processor and (b) a computer-readable recording medium of the present invention coupled to said computer processor.
[0378] The recording medium, device, and computer program of the present invention enable the method of the present invention described above to be carried out on a computer, and common details among them are omitted to avoid excessive complexity in this specification due to repetitive descriptions.
[0379] Program instructions, when executed by a processor, cause the processor to execute the method of the present invention described above. Program instructions for executing a method for providing a nucleic acid sequence data set for designing oligonucleotides used to detect a target nucleic acid molecule of an organism of interest may include the following instructions: (i) an instruction to receive the name of the target nucleic acid molecule and the name of the organism of interest, and to retrieving synonyms for the target nucleic acid molecule of the organism of interest; (ii) an instruction to collect nucleic acid sequence data contained in nucleic acid records relating to the organism of interest, wherein at least one of the name of the target nucleic acid molecule and the collected synonyms is listed in a descriptor; (iii) an instruction to sort the collected nucleic acid sequence data according to a taxonomic name and / or taxonomic ID and to select a taxonomic representative sequence from among the nucleic acid sequence data having the same taxonomic name and / or taxonomic ID; (iv) an instruction to group the selected taxonomic representative sequences according to homology and to select a group representative sequence from each group; and (v) an instruction to collect nucleic acid sequence data having homology greater than a predetermined value with the group representative sequence and provide it as a nucleic acid sequence data set for designing the oligonucleotide (e.g., to display on an output device).
[0380] The method of the present invention is executed in a processor, and the processor may be a processor in a data acquisition device such as a stand-alone computer, a network-attached computer, or a real-time PCR device.
[0381] Computer-readable recording media include, but are not limited to, various storage media known in the art, such as CD-R, CD-ROM, DVD, flash memory, floppy disk, hard drive, portable HDD, USB, magnetic tape, MINIDISC, non-volatile memory card, EEPROM, optical disc, optical storage media, RAM, ROM, system memory, and web server.
[0382] Nucleic acid sequence data sets for designing oligonucleotides used to detect target nucleic acid molecules of an organism of interest can be provided in various ways. For example, nucleic acid sequence data sets for designing oligonucleotides can be provided to a separate system, such as a desktop computer system, via a network connection (e.g., LAN, VPN, Internet, and intranet) or a direct connection (e.g., USB or other direct wired or wireless connection), or can be provided on portable media such as CDs, DVDs, floppy disks, and portable HDDs. Similarly, nucleic acid sequence data sets for designing oligonucleotides can be provided to a server system via a network connection (e.g., LAN, VPN, Internet, intranet, and wireless communication networks) to a client, such as a laptop or desktop computer system.
[0383] Instructions for implementing a processor that executes the present invention may be included in a logic system. Although said instructions may be provided on a software recording medium (e.g., portable HDD, USB, floppy disk, CD, and DVD), they may be downloadable and stored in a memory module (e.g., a hard drive or other memory such as local or attached RAM or ROM). Computer code that executes the present invention may be executed in various coding languages such as C, C++, Java, Visual Basic, VBScript, JavaScript, Perl, and XML. Additionally, various languages and protocols may be used for the external and internal storage and transmission of signals and commands according to the present invention.
[0384] A computer processor can be constructed so that a single processor performs all of the aforementioned performances. Alternatively, a processor unit can be constructed so that multiple processors execute each performance. Effects of the invention
[0385] The features and advantages of the present invention are summarized as follows:
[0386] (a) A conventional method for providing a nucleic acid sequence data set for designing oligonucleotides used to detect target nucleic acid molecules of an organism of interest involves collecting nucleic acid sequence data using keywords such as the name of the target nucleic acid molecule, aligning the collected nucleic acid sequence data according to sequence length and determining the longest sequence as the representative sequence, grouping nucleic acid sequence data having homology greater than a predetermined value with the representative sequence, and then integrating the nucleic acid sequence data collected with the keywords with the nucleic acid sequence data having homology with the representative sequence to provide a target nucleic acid sequence data set for the target nucleic acid molecule. Additionally, an alignment file for the target nucleic acid sequence data set is provided for use in designing oligonucleotides.
[0387] As a result, it was confirmed that proper alignment was not achieved due to differences in homology between the sequences collected from the aforementioned representative sequences, leading to a problem where analysts spent unnecessary time reviewing the aligned nucleic acid sequences.
[0388] (b) To solve the aforementioned problem, the present invention sorts nucleic acid sequence data collected from synonyms for target nucleic acid molecules according to taxonomic names and / or taxonomic IDs, selects a taxonomic representative sequence from among the nucleic acid sequence data having the same taxonomic name and / or taxonomic ID, groups the selected taxonomic representative sequences according to homology, and selects a group representative sequence from each group, thereby providing nucleic acid sequence data having homology of a predetermined value or more with the group representative sequence.
[0389] As a result, it was confirmed that multiple target nucleic acid sequences for the above-mentioned target nucleic acid molecule were collected without omission, and that the alignment results of the collected multiple target nucleic acid sequences were properly formed so that they could be referenced in the design of the oligonucleotide.
[0390] (c) According to the present invention, the alignment result of the nucleic acid sequence data set is properly formed so that it can be used for the design of oligonucleotides, thereby solving the time-consuming and labor-consuming problems of analysts reviewing errors in the entry of sequences included in the collected nucleic acid sequence data set. Brief explanation of the drawing
[0391] Figure 1 is a flowchart showing the process of providing a target nucleic acid sequence data set according to a conventional method (International Publication No. WO2019 / 212238). Figure 2 shows Salmonella enterica (according to the above conventional method) Salmonella entericaShows the alignment results of the nucleic acid sequence dataset for the design of oligonucleotides used to detect the gene sopB of ). FIG. 3 is a flowchart showing the process of providing a nucleic acid sequence data set for designing oligonucleotides used to detect target nucleic acid molecules of an organism of interest according to one embodiment of the present invention. Figure 4 shows the name of the target nucleic acid molecule (ompA) and the name of the organism of interest (in the UI (User Interface) Chlamydophila pneumoniae As a result of inputting ), gene information summary records are collected from the gene database of NCBI (National Center for Biotechnology Information) or from a gene database built by downloading the said gene database, and protein names are collected from gene descriptions as synonyms for target nucleic acid molecules in said records. Figure 5 shows the organism of interest Enterobacter cloacae complex If the target nucleic acid molecule is ompX, it shows the collection of identifiers of nucleic acid records. Looking at Figure 5, Accession No and GI No can be identified as identifiers. Figure 6 is a captured image of a portion of the nucleic acid record that appears when the title of Accession: CP017990.1 in Figure 5 is clicked. In the above nucleic acid record gene , CDS, / gene, / note, / product, etc. represent descriptors. Figure 7 shows the organism of interest ( Enterobacter cloacaeThis document describes the process of collecting nucleic acid sequence data identified by collected identifiers using the name (ompX) and synonym (outer membrane protein) of the target nucleic acid molecule of the complex, aligning the nucleic acid sequence data according to alignment criteria among those with the same taxonomic identifier, and selecting a taxonomic representative sequence. FIG. 8 shows Salmonella enterica (according to one embodiment of the present invention) Salmonella enterica Displays a User Interface (UI) with sopB entered as the name of the target nucleic acid molecule for ). Salmonella enterica as the organism of interest ( Salmonella enterica ) and its taxonomic identifier (Taxonomic ID: 28901) are entered by clicking In / Exclusivity in the UI above. FIG. 9 shows Salmonella enterica (according to one embodiment of the present invention) Salmonella enterica Displays a User Interface (UI) with sopB entered as the name of the target nucleic acid molecule for ) and the protein name (inositol phosphatase) as a synonym. Salmonella enterica as the organism of interest ( Salmonella enterica ) and its taxonomic identifier (Taxonomic ID: 28901) are entered by clicking In / Exclusivity in the UI above. FIG. 10 is an organism of interest provided according to one embodiment of the present invention ( Salmonella enterica Shows the alignment results of the target nucleic acid sequence dataset for the target nucleic acid molecule (sopB) of ). FIG. 11 is an organism of interest provided according to another embodiment of the present invention ( Salmonella enterica Shows the alignment results of the target nucleic acid sequence dataset for the target nucleic acid molecule (sopB) of ). FIG. 12 shows an organism of interest provided according to a conventional method (International Publication No. WO2019 / 212238) ( Salmonella entericaAfter the analyst reviews the alignment results of the target nucleic acid sequence dataset for the target nucleic acid molecule (sopB) of ), the alignment results provided after running the program 4 times are displayed. Specific details for implementing the invention
[0392] The present invention will be described in more detail below through examples. These embodiments are intended solely to explain the present invention more specifically, and it will be obvious to those skilled in the art that the scope of the present invention is not limited by these embodiments according to the gist of the present invention.
[0393] Examples
[0394] Example 1: Salmonella enterica ( Salmonella enterica Provision of a nucleic acid sequence data set for the design of oligonucleotides used to detect the gene sopB of ).
[0395] A program (AutoMSA v3.0) that provides nucleic acid sequence data sets for designing oligonucleotides used to detect target nucleic acid molecules of organisms of interest was executed to provide nucleic acid sequence data sets for designing oligonucleotides used to detect the gene sopB of Salmonella enterica.
[0396] The scientific name of Salmonella enterica ( Salmonella enterica The ) and gene name (sopB) were entered into the User Interface (UI) window of the AutoMSA v3.0 program (Fig. 8) and the AutoMSA v3.0 program was executed.
[0397] The AutoMSA v3.0 program was executed and the following steps were followed: (1) gene name (sopB) and Salmonella enterica ( Salmonella enterica) was received, and synonyms for the sopB of the above Salmonella enterica (protein names: inositol phosphatase or inositol phosphate phosphatase, etc.) were collected from a gene database (using a gene database constructed by downloading the NCBI gene database). Specifically, when sopB was listed as a gene symbol in the Full report of the gene information summary record, inositol phosphatase or inositol phosphate phosphatase, etc. listed in the gene description were collected as synonyms. In addition, inositol phosphatase or inositol phosphate phosphatase, etc. listed in the Summary of the gene information summary record were collected as synonyms.
[0398] Next, the above Salmonella enterica, the gene name (sopB), and a synonym for sopB (protein name) were received, and a set of nucleic acid sequence data was collected relating to the above Salmonella enterica, which includes nucleic acid records in which the sopB and / or the synonym is listed in the descriptor of the nucleic acid record.
[0399] (2) The above Salmonella enterica, the gene name (sopB), and a synonym for sopB (protein name) were received, and the identifiers of nucleic acid records relating to the above Salmonella enterica, in which the sopB and / or the synonym are listed in the descriptor of the nucleic acid record, were collected from a nucleotide database (using the nucleotide database of NCBI, i.e., a nucleotide database constructed by downloading the nucleotide database including NCBI’s GenBank (including STS, EST, GSS, SNP, TSA, PAT, WGS and non-WGS databases), RefSeq, DDBJ and EMBL databases). Specifically, the received Salmonella enterica, sopB, inositol phosphatase or inositol phosphate phosphatase, etc. were entered as a query in a nucleotide database, and the organism name (taxonomic name), taxonomic identifier (Taxonomic ID), sopB, inositol phosphatase or inositol phosphate phosphatase, etc. were entered as descriptors of nucleic acid records in title, gene, CDS, / gene, / product, / note or / taxon, etc., and the identifiers (specifically, access number or GI, etc.) of nucleic acid records in which sopB, inositol phosphatase or inositol phosphate phosphatase, etc. were listed. (3) Nucleic acid sequence data identified by the above identifiers were collected. Specifically, nucleic acid records identified by the above identifiers were collected, and nucleic acid sequence data was collected from nucleic acid records that are identical to or contain sopB, inositol phosphatase or inositol phosphate phosphatase in gene, CDS, / gene, / product or / note as descriptors of the above nucleic acid records.
[0400] (4) The collected nucleic acid sequence data were sorted according to taxonomic names and / or taxonomic IDs, and a representative taxonomic sequence was selected from among the nucleic acid sequence data having the same taxonomic name and / or taxonomic ID. Specifically, the selection of the representative taxonomic sequence from among the nucleic acid sequence data having the same taxonomic name and / or taxonomic ID was carried out as follows: First, the nucleic acid sequence data having the same taxonomic name and / or taxonomic ID were sorted according to the sorting criteria having the following order:
[0401] (i) sorted according to the assembly level of the nucleic acid sequences included in the above nucleic acid sequence data; the assembly levels are ranked in the order of complete genome, chromosome, scaffold, and contig (i.e., the complete genome has the highest rank),
[0402] (ii) Sort according to whether the above nucleic acid sequence data is included in the RefSeq (Reference Sequence) database; if the above nucleic acid sequence data is included in the RefSeq database, it is ranked higher than if it is not included, and
[0403] (iii) sorting based on whether the name of the nucleic acid molecule listed in the descriptor of the nucleic acid record containing the nucleic acid sequence data matches the name of the received target nucleic acid molecule and at least one of the collected synonyms; the case of a match is ranked higher than the case of a non-match, and
[0404] (iv) sorted according to the length of the nucleic acid sequence data; the longer the length, the higher the rank, and
[0405] (v) Sort according to whether a host is listed in the descriptor of the nucleic acid record containing the above nucleic acid sequence data; the case where the host of interest for the organism of interest is listed in the host is higher than the case where it is not listed, and the case where it is not listed is higher than the case where an organism other than the host of interest is listed in the host, and
[0406] (vi) Sort according to the registration date or modification date of the nucleic acid record containing the above nucleic acid sequence data; the newer the registration date or modification date, the higher the rank, and
[0407] (vii) Sort according to the alphabet of the access number of the above nucleic acid sequence data; the earlier the alphabet of the access number, the higher the rank.
[0408] And, among the above-mentioned aligned nucleic acid sequence data, the nucleic acid sequence of the top-ranked nucleic acid sequence data was selected as the taxonomic representative sequence. In the above alignment criteria (v), the host of interest is Homo sapiens ( Homo sapiens )am.
[0409] This process was performed on nucleic acid sequence data aligned by taxonomic name and / or taxonomic identifier.
[0410] (5) The selected taxonomic representative sequences were grouped according to homology, and a group representative sequence was selected from each group. Specifically, the selected taxonomic representative sequences were aligned according to the following alignment criteria:
[0411] (i) Aligned according to the assembly level of the selected taxonomic representative sequences; the assembly levels are ranked in the order of complete genome, chromosome, scaffold, and contig (i.e., the complete genome has the highest rank),
[0412] (ii) the number of nucleic acid sequence data having the same taxonomic name and / or taxonomic identification symbol as the selected taxonomic representative sequence; the higher the number, the higher the rank, and
[0413] (iii) Sort according to whether a host is listed in the descriptor of the nucleic acid record containing the selected taxonomic representative sequence; the case where the host of interest for the organism of interest is listed in the host is higher in rank than the case where it is not listed, and the case where it is not listed is higher in rank than the case where an organism other than the host of interest is listed in the host, and
[0414] (iv) Sort according to the alphabet of the accession number of the selected taxonomic representative sequence; the earlier the alphabet of the accession number, the higher the rank. In the sorting criteria (iii), the host of interest is Homo sapiens ( Homo sapiens )am.
[0415] Among the above-mentioned aligned taxonomic representative sequences, the highest taxonomic representative sequence was selected, and taxonomic representative sequences having 90% or more homology with the above-mentioned highest taxonomic representative sequence were grouped using the UCLUST algorithm, and the above-mentioned highest taxonomic representative sequence in each group was selected as the group representative sequence.
[0416] (6) Nucleic acid sequence data having 50% or more homology with the above group representative sequence (specifically, Identity: 50% or more, word size: 15, E-value: 10000) was collected from a nucleotide database (specifically, using the nucleotide database of (2) above) by performing BLAST and provided as a nucleic acid sequence data set for designing the above oligonucleotide. (7) Among the provided nucleic acid sequence data sets, the received nucleic acid sequence data regarding Salmonella enterica was provided as a target nucleic acid sequence data set for the gene sopB. The provided target nucleic acid sequence data set has 90% or more homology (Identity: 90% or more) with at least one representative sequence among the group representative sequence and the taxonomic representative sequence for the target nucleic acid sequence data set. The target nucleic acid sequence data set having 90% or more homology is provided as follows. Specifically, the sequences of the target nucleic acid sequence dataset were extended to the length of the group representative sequence and / or taxonomic representative sequence, and then sequences with a query coverage of 10% or more and an identity of 90% or more were selected based on the length of the group representative sequence and / or taxonomic representative sequence.
[0417] (8) Among the nucleic acid sequence data sets provided in (6) above, the nucleic acid sequence data that is not related to the received Salmonella enterica was provided as a non-target nucleic acid sequence data set for non-target nucleic acid molecules. And, the provided non-target nucleic acid sequence data set has (i) 100% homology (Identity 100%) for a 20 bp sequence region of at least one representative sequence among the group representative sequence and the taxonomic representative sequence, and (ii) 70% or more homology (Identity 70% or more) for at least one representative sequence among the group representative sequence and the taxonomic representative sequence.
[0418] Specifically, (i) the sequences of the non-target nucleic acid sequence dataset are extended by the length of the group representative sequence and / or taxonomic representative sequence, and then a non-target nucleic acid sequence dataset is selected that has 100% homology (Identity 100%) for a 20 bp sequence region while moving a 20 bp sequence region from one end of at least one representative sequence and / or the sequences of the non-target nucleic acid sequence dataset, and (ii) among the selected non-target nucleic acid sequence datasets, sequences with Query coverage 100% and Identity 70% or more are selected based on the length of the group representative sequence and / or taxonomic representative sequence.
[0419] Figure 10 shows the alignment results of the target nucleic acid sequence dataset for the gene sopB of Salmonella enterica provided in (7) above. As can be seen in Figure 10, it can be confirmed that the sequences were properly aligned according to homology as a result of aligning multiple target nucleic acid sequences. In the alignment results of Figure 10, it was determined that the more black shading there is compared to gray shading, the more properly the alignment was formed according to homology. There were 11 representative sequences of the group selected in (5) above.
[0420] As a result of executing the AutoMSA v3.0 program according to Example 1, a list of nucleic acid sequence data sets for designing oligonucleotides of (6), a list of target nucleic acid sequence data sets of (7), and a list of non-target nucleic acid sequence data sets of (8) are provided, which include information such as access number, collected database information, length of nucleic acid sequence, location information of the gene, whether it is a taxonomic representative sequence or a group representative sequence, organism name (taxonomic name), taxonomic identifier, homology information, bio sample number, assembly level, and RefSeq number. Additionally, an alignment file of the target nucleic acid sequence data set of (7) and an alignment file of the non-target nucleic acid sequence data set of (8) are provided.
[0421] Comparative Example 1: Salmonella enterica ( Salmonella enterica Provision of a nucleic acid sequence data set for the design of oligonucleotides used to detect the gene sopB of ).
[0422] The nucleic acid sequence data set for the design of the oligonucleotide used to detect the gene sopB of Salmonella enterica provided in Comparative Example 1 was provided as an AutoMSA program (AutoMSA v2.0) according to the method described in International Publication No. WO2019 / 212238 filed by the applicant.
[0423] The conventional AutoMSA program (AutoMSA v2.0) is identical to the program sequence (1) to (3) of Example 1, but differs from the program sequence (4) to (8) of Example 1 as follows. The parts that differ from Example 1 are explained as follows.
[0424] The scientific name of Salmonella enterica ( Salmonella enterica ) and gene name (sopB) were entered into the User Interface (UI) window of the above AutoMSA v2.0 program, and the AutoMSA v2.0 program was executed.
[0425] The AutoMSA v2.0 program was executed and the process was carried out in the following order: (1) to (3) were carried out in the same manner as in Example 1. (4) The collected nucleic acid sequence data were sorted according to sequence length, the longest nucleic acid sequence data among the sorted nucleic acid sequence data was selected, and nucleic acid sequence data having 90% or more homology (Identity: 90% or more) with the longest nucleic acid sequence data was grouped using the UCLUST algorithm, and the longest nucleic acid sequence data in each group was selected as the group representative sequence. (5) This was carried out in the same manner as in (6) of Example 1. (6) A target nucleic acid sequence data set was provided by carrying out the process in the same manner as in (7) of Example 1, excluding the content regarding the taxonomic representative sequence in (7) of Example 1.
[0426] However, according to the conventional AutoMSA v2.0 program, the nucleic acid sequence data set collected in (3) above and the target nucleic acid sequence data set provided in (6) above were integrated to provide a target nucleic acid sequence data set for the design of oligonucleotides.
[0427] The results of aligning the target nucleic acid sequence dataset for oligonucleotide design provided according to the conventional AutoMSA v2.0 program are shown in Figure 2. As can be seen in Figure 2, the presence of many gray shades indicates that the alignment of multiple nucleic acid sequences was not properly formed according to homology. Upon investigating the reason for this failure to properly form the alignment, it was found that there were 25 group representative sequences, and due to differences in sequence homology between these group representative sequences, differences in homology also occurred among the sequences collected as said group representative sequences, and as a result, the alignment of multiple nucleic acid sequences was not properly formed.
[0428] Example 2: Salmonella enterica ( Salmonella enterica Provision of a nucleic acid sequence data set for the design of oligonucleotides used to detect the gene sopB of ).
[0429] A nucleic acid sequence data set for designing oligonucleotides used to detect the sopB gene of Salmonella enterica was provided by executing an AutoMSA v3.0 program that was carried out in a manner similar to Example 1, but with the addition of an algorithm that automatically performs the process of removing duplicate sequences and reviewing sequences with errors in the nucleotide database in the program.
[0430] First, the protein name of the gene sopB was searched in the NCBI gene database and confirmed to be inositol phosphatase.
[0431] In the User Interface (UI) window of the AutoMSA v3.0 program, the scientific name of Salmonella enterica ( Salmonella enterica ), gene name (sopB) and protein name (inositol phosphatase) were entered (Fig. 9) and the AutoMSA v3.0 program was executed.
[0432] The AutoMSA v3.0 program was executed and the following steps were performed: (1) Salmonella enterica ( Salmonella enterica), received the gene name (sopB) and the protein name (inositol phosphatase), and collected synonyms for the sopB of the above Salmonella enterica (protein name: inositol phosphate phosphatase, etc.) from a gene database (identical to the gene database used in (1) of Example 1 above), and after reviewing the protein name inositol phosphate phosphatase, etc., collected as a synonym for the above sopB, inositol phosphate phosphatase, etc. was used as a synonym. Specifically, when sopB was listed as a gene symbol in the Full report of the gene information summary record, inositol phosphatase or inositol phosphate phosphatase, etc. listed in the gene description were collected as a synonym. Then, inositol phosphatase or inositol phosphate phosphatase, etc. listed in the Summary of the gene information summary record were collected as a synonym. Here, since inositol phosphatase is a synonym entered by the user before the program execution, inositol phosphate phosphatase, etc. are additionally collected synonyms.
[0433] Next, the above Salmonella enterica, the gene name (sopB), and a synonym for sopB (protein name) were received, and nucleic acid sequence data contained in nucleic acid records relating to the above Salmonella enterica, in which the sopB and / or the synonym (protein name) is listed in the descriptor of the nucleic acid record were collected.
[0434] (2) The above Salmonella enterica, the gene name (sopB), and a synonym for sopB (protein name) were received, and identifiers of nucleic acid records relating to the above Salmonella enterica, in which the sopB and / or the synonym (protein name) are listed in the descriptor of the nucleic acid record, were collected from a nucleotide database. Specifically, the received Salmonella enterica,sopB, inositol phosphatase, or inositol phosphate phosphatase, etc. were entered as a query into a nucleotide database (identical to the nucleotide database of (3) in Example 1 above), and the organism name (taxonomic name), taxonomic identifier (Taxonomic ID), sopB, inositol phosphatase, or inositol phosphate phosphatase, etc. were listed as descriptors of nucleic acid records such as title, gene, CDS, / gene, / product, / note, or / taxon, etc. (specifically, access number or GI, etc.) of the nucleic acid records. (3) Nucleic acid sequence data identified by the above identifiers were collected. Specifically, nucleic acid records identified by the above-mentioned identifiers were collected, and nucleic acid sequence data was collected from nucleic acid records in which gene, CDS, or / gene is identical to or contains sopB as a descriptor of the nucleic acid records, and inositol phosphatase or inositol phosphate phosphatase is identical to or contains inositol phosphatase as a descriptor of / product. Where nucleic acid sequence data was not collected through descriptors for gene names and protein names, nucleic acid sequence data was collected from nucleic acid records in which / note is identical to or contains sopB, inositol phosphatase, or inositol phosphate phosphatase.
[0435] (3-1) Duplicate sequences were removed from the collected nucleic acid sequence data. Specifically, the collected nucleic acid sequence data were sorted according to biosample identifiers to select nucleic acid sequence data having the same biosample identifier, and the selected nucleic acid sequence data were sorted according to the following sorting criteria: (i) sorted the selected nucleic acid sequence data according to assembly level; the assembly levels are ranked in the order of complete genome, chromosome, scaffold, and contig (i.e., the complete genome has the highest rank), and (ii) sorted according to whether the selected nucleic acid sequence data is included in the RefSeq (Reference Sequence) database; the nucleic acid sequence data has a higher rank when it is included in the RefSeq database than when it is not included.
[0436] Then, the nucleic acid sequence of the top-most nucleic acid sequence data among the above-mentioned aligned nucleic acid sequence data was selected, and nucleic acid sequence data other than the top-most nucleic acid sequence data was removed from the above-mentioned collected nucleic acid sequence data.
[0437] (4 to 8) The above-mentioned method was carried out in the same manner as (4) to (8) of Example 1. However, the alignment criterion (iii) of (4) of Example 1 was carried out according to the following alignment criterion: (iii) alignment according to whether the name of the nucleic acid molecule and at least one of the collected synonyms and the name of the nucleic acid molecule and / or the name of the protein listed in the descriptor of the nucleic acid record containing the nucleic acid sequence data match in the order of the name of the nucleic acid molecule and the name of the protein, and the name of the nucleic acid molecule; if they match in the above order, the rank is high (i.e., the rank is highest when the name of the nucleic acid molecule and at least one of the collected synonyms and the name of the nucleic acid molecule and the name of the protein listed in the descriptor of the nucleic acid record containing the nucleic acid sequence data all match), and if they do not match, the rank is lowest.
[0438] (9) An error in the registration of the group representative sequence included in the target nucleic acid sequence data set provided in (7) above was checked. Specifically, the process was carried out by a method including the following steps: nucleic acid sequence data having 90% or more homology (Identity: 90% or more) with the group representative sequence was collected from a nucleotide database (specifically, using a nucleotide database built by downloading a nucleotide database including NCBI’s GenBank (including SNP and non-WGS databases), RefSeq, DDBJ, and EMBL databases) to provide a nucleic acid sequence data set (9-1) (in addition, nucleic acid sequence data for Salmonella enterica among the nucleic acid sequence data sets of (6), nucleic acid sequence data having 90% or more homology with the group representative sequence, and then a nucleic acid sequence data set included in the nucleotide database of (9) may be selected and provided), and if the taxonomic name and / or taxonomic identification symbol of the organism for the group representative sequence satisfies one of the following predetermined criteria, the target nucleic acid sequence data of the group representative sequence and the same group as the group representative sequence in the target nucleic acid sequence data set of (7) The target nucleic acid sequence data belonging to the group was classified as design-excluded target nucleic acid sequence data. And, the above-mentioned predetermined criteria include the following: (i) where there is no nucleic acid sequence data corresponding to the taxonomic name and / or taxonomic identification symbol of the organism for the group representative sequence in the nucleic acid sequence data set of (9-1), and only nucleic acid sequence data corresponding to the taxonomic name and / or taxonomic identification symbol of an organism other than the organism for the group representative sequence exists;
[0439] (ii) If the homology of the target nucleic acid sequence data corresponding to the taxonomic name and / or taxonomic identification symbol of the organism for the group representative sequence in the nucleic acid sequence data set of (9-1) above is lower compared to the homology of the nucleic acid sequence data corresponding to the taxonomic name and / or taxonomic identification symbol of the organism and other organisms for the group representative sequence;
[0440] (iii) in which there is no target nucleic acid sequence data corresponding to the taxonomic name and / or taxonomic identification symbol of an organism for the group representative sequence in the nucleic acid sequence data set of (9-1), and the ratio of nucleic acid sequence data corresponding to an organism corresponding to a taxonomic name and / or taxonomic identification symbol higher than that of the organism for the group representative sequence or an organism corresponding to a taxonomic name and / or taxonomic identification symbol lower than that thereof is less than 10% with respect to the nucleic acid sequence data set of (9-1); and
[0441] (iv) where all of the nucleic acid sequence data sets of (9-1) above are nucleic acid sequence data sets of organisms for the group representative sequences, but the name of the target nucleic acid molecule is not listed in the descriptors of the nucleic acid records containing the nucleic acid sequence data sets, or is different from at least one of the name of the target nucleic acid molecule and the collected synonyms.
[0442] In the above (9), the group representative sequence satisfying the above-mentioned criteria is a group representative sequence with an entry error. Therefore, in the target nucleic acid sequence data set provided in the above (7), the group representative sequence with the entry error and the target nucleic acid sequence data set belonging to the same group as the group representative sequence are classified as design-excluded target nucleic acid sequence data sets that are not used in the design of oligonucleotides.
[0443] (10) The process of removing duplicate sequences of (3-1) above was carried out after (6) or (9) above.
[0444] (11) It was checked whether there were any errors in the non-target nucleic acid sequence data set provided in (8) above. Specifically, this was carried out by a method including the following process: Nucleic acid sequence data having 90% or more homology with the non-target nucleic acid sequence of the organism in the non-target nucleic acid sequence data set was collected from a nucleotide database (specifically, a nucleotide database built by downloading a nucleotide database including NCBI’s GenBank (including SNP and non-WGS databases), RefSeq, DDBJ and EMBL databases) to provide a nucleic acid sequence data set (11-1), and if the taxonomic name and / or taxonomic identification symbol of the organism in the non-target nucleic acid sequence satisfied one of the following criteria, the non-target nucleic acid sequence data of the organism in the non-target nucleic acid sequence data set was classified as design-excluded non-target nucleic acid sequence data. And, the above-mentioned predetermined criteria include the following: (i) where the nucleic acid sequence data set of (11-1) above does not contain nucleic acid sequence data corresponding to the taxonomic name and / or taxonomic identification symbol of the organism for the non-target nucleic acid sequence, and only nucleic acid sequence data corresponding to the taxonomic name and / or taxonomic identification symbol of the organism and other organisms exists;
[0445] (ii) if the homology of the non-target nucleic acid sequence data corresponding to the taxonomic name and / or taxonomic identification symbol of the organism for the non-target nucleic acid sequence in the nucleic acid sequence data set of (11-1) is lower compared to the homology of the nucleic acid sequence data corresponding to the taxonomic name and / or taxonomic identification symbol of the organism and other organisms; and
[0446] (iii) In the nucleic acid sequence data set of (11-1) above, there is no non-target nucleic acid sequence data corresponding to the taxonomic name and / or taxonomic identification symbol of the organism for the non-target nucleic acid sequence, and the ratio of nucleic acid sequence data corresponding to an organism corresponding to a taxonomic name and / or taxonomic identification symbol higher than the organism or an organism corresponding to a taxonomic name and / or taxonomic identification symbol lower than the organism is less than 10% of the nucleic acid sequence data set of (11-1).
[0447] In the above (11), the organism for the non-target nucleic acid sequence data set satisfying the above-mentioned criteria is a non-target nucleic acid sequence data set with an entry error, so the non-target nucleic acid sequence data set of the organism with the entry error in the non-target nucleic acid sequence data set provided in the above (8) was classified as a design-excluded non-target nucleic acid sequence data set that is not used in the design of oligonucleotides.
[0448] FIG. 11 shows the alignment results of the target nucleic acid sequence dataset for the gene sopB of Salmonella enterica provided after (10). As can be seen in FIG. 11, the alignment of multiple target nucleic acid sequences resulted in more black shading than gray shading, confirming that the sequences were properly aligned according to homology. There were 5 group representative sequences selected in (10), and the total number of taxonomic names or taxonomic identifiers included in the final collected target nucleic acid sequence dataset was 1,549, and the number of final collected target nucleic acid sequence data was 13,989.
[0449] As a result of comparing FIG. 10, which is the alignment result provided by Example 1 above, with FIG. 11, which is the alignment result provided by Example 2 above, it was confirmed that the alignment of multiple target nucleic acid sequences was better formed in FIG. 11.
[0450] As a result of executing the AutoMSA v3.0 program according to Example 2, a list of nucleic acid sequence data sets for the design of oligonucleotides of (6) containing information such as access number, collected database information, length of nucleic acid sequence, location information of genes, whether it is a taxonomic representative sequence or a group representative sequence, organism name (taxonomic name), taxonomic identifier, homology information, bio sample number, assembly level, RefSeq number, etc., a list of target nucleic acid sequence data sets of (7), a list of non-target nucleic acid sequence data sets of (8), a list of nucleic acid sequence data sets removed as duplicate sequences in (3-1) and (10), a list of design-excluded target nucleic acid sequence data sets classified in (9), and a list of design-excluded non-target nucleic acid sequence data sets classified in (11) were provided. Additionally, an alignment file of the target nucleic acid sequence data set of (7) and an alignment file of the non-target nucleic acid sequence data set of (8) were provided.
[0451] Comparative Example 2: Salmonella enterica ( Salmonella enterica Provision of a nucleic acid sequence data set for the design of oligonucleotides used to detect the gene sopB of ).
[0452] In Comparative Example 1, the user who received the alignment file of Fig. 2 provided by the AutoMSA program (AutoMSA v2.0) was unable to design oligonucleotides from the alignment results of Fig. 2, so the user performed a process of reviewing group representative sequences with registration errors among the aligned nucleic acid sequences and deleting group representative sequences that break the alignment between multiple nucleic acid sequences. When the user performs a review process such as deleting group representative sequences, the user must re-collect nucleic acid sequence data having homology to the changed group representative sequences from the nucleotide database, so the AutoMSA program (AutoMSA v2.0) was restarted.
[0453] Among the 25 group representative sequences selected in Comparative Example 1, group representative sequences 1, 3, and 7-24, which had user-entered errors, and the target nucleic acid sequence data set belonging to the same group as the group representative sequences were deleted four times, and the AutoMSA program (AutoMSA v2.0) was restarted four times, resulting in an alignment result as shown in Fig. 12.
[0454] In Comparative Example 2 above, there were 5 group representative sequences selected through the user's 4 sequence review processes, and the total number of taxonomic names or taxonomic identifiers included in the final collected target nucleic acid sequence data set was 1,517, and the number of final collected target nucleic acid sequence data was 13,798.
[0455] When comparing the results of Example 2 and Comparative Example 2, the automated sequence collection method (AutoMSA v3.0) according to Example 2 provided an alignment file sufficient to design oligonucleotides by selecting five representative group sequences with a single program run, and collected more target nucleic acid sequences than Comparative Example 2.
[0456] Specific parts of the present invention have been described in detail above. It is evident to those skilled in the art that such specific descriptions are merely preferred embodiments and do not limit the scope of the invention. Accordingly, the actual scope of the invention is defined by the appended claims and their equivalents.
Claims
Claim 1 A computer-implemented method for providing a nucleic acid sequence data set for designing oligonucleotides used to detect a target nucleic acid molecule of an organism of interest, comprising the following steps: (a) receiving the name of the target nucleic acid molecule and the name of the organism of interest, and retrieving synonyms for the target nucleic acid molecule of the organism of interest; (b) collecting nucleic acid sequence data contained in nucleic acid records; Each of the above nucleic acid records relates to the organism of interest and includes a descriptor in which the name of the target nucleic acid molecule and at least one of the collected synonyms are listed; (c) sorting the collected nucleic acid sequence data according to a taxonomic name and / or a taxonomic ID and selecting a taxonomic representative sequence from among the nucleic acid sequence data having the same taxonomic name and / or taxonomic ID; (d) grouping the selected taxonomic representative sequences according to homology and selecting a group representative sequence from each group; and (e) collecting nucleic acid sequence data having homology of a predetermined value or more with the group representative sequence and providing it as a nucleic acid sequence data set for designing the oligonucleotide. Claim 2 A method according to claim 1, wherein step (a) is a step of receiving the name of a target nucleic acid molecule, the name of a protein encoded by the target nucleic acid molecule, and the name of an organism, and retrieving synonyms of the organism for the target nucleic acid molecule and the protein. Claim 3 A method according to claim 1, wherein the nucleic acid sequence data of step (b) comprises nucleic acid sequence data corresponding to part or all of the target nucleic acid molecule or variant nucleic acid sequence data for the target nucleic acid molecule. Claim 4 A method according to claim 1, wherein the collection of nucleic acid sequence data of step (b) is carried out by a method comprising the following steps: (b-1) collecting identifiers of nucleic acid records; each of the nucleic acid records relates to the organism of interest and includes a descriptor in which at least one of the name of the target nucleic acid molecule and the collected synonyms is listed; and (b-2) collecting nucleic acid sequence data specified by the identifiers. Claim 5 A method according to claim 4, wherein step (b-1) relates to the organism of interest and is characterized by collecting identifiers of nucleic acid records in which at least one of the name of the target nucleic acid molecule, the name of the protein, and the collected synonyms is listed in the descriptor. Claim 6 A method according to claim 4, wherein step (b-2) is characterized by selectively collecting nucleic acid sequence data corresponding to the target nucleic acid molecule within the nucleic acid sequence data specified by the identifiers. Claim 7 A method according to claim 4, wherein step (b-2) comprises the following steps: (b-2-1) collecting nucleic acid records identified by the identifiers; and (b-2-2) collecting nucleic acid sequence data corresponding to a target nucleic acid molecule from the nucleic acid records. Claim 8 In claim 7, the step (b-2-2) is characterized by selectively collecting nucleic acid sequence data corresponding to a target nucleic acid molecule and identification information of said nucleic acid sequence data from said nucleic acid records, and the step of selectively collecting said nucleic acid sequence data and identification information of said nucleic acid sequence data comprises the following steps: (b-2-2-1) determining, among one or more sub-records within said nucleic acid records, a sub-record in which said synonym is recorded in a predetermined first specification as a valid sub-record; (b-2-2-2) if there is no valid sub-record determined by the first specification within said nucleic acid records, a sub-record in which said synonym is recorded in a second specification as a valid sub-record; (b-2-2-3) if there is no valid sub-record determined by the second specification within said nucleic acid records, a sub-record in which said synonym is recorded in a third specification as a valid sub-record; and (b-2-2-4) a step of collecting nucleic acid sequence data and identification information for the valid sub-records determined above. Claim 9 A method according to claim 1, characterized by additionally including the following steps between steps (b) and (c): (b-3) sorting the collected nucleic acid sequence data according to a biosample identifier to select nucleic acid sequence data having the same biosample identifier; (b-4) sorting the selected nucleic acid sequence data to satisfy at least one of the following sorting criteria; (b-5) selecting the nucleic acid sequence of the top nucleic acid sequence data among the sorted nucleic acid sequence data; and (b-6) removing nucleic acid sequence data other than the top nucleic acid sequence data from the collected nucleic acid sequence data, wherein the sorting criteria include: (i) sorting the selected nucleic acid sequence data according to the assembly level; The assembly levels are ranked in the order of complete genome, chromosome, scaffold, and contig, with the highest ranking; (ii) sort the selected nucleic acid sequence data according to whether the selected nucleic acid sequence data is included in the RefSeq (Reference Sequence) database; the ranking is higher when the nucleic acid sequence data is included in the RefSeq database than when it is not included. Claim 10 A method according to claim 1, wherein the selection of the taxonomic representative sequence in step (c) is carried out by a method comprising the following steps: (c-1) aligning nucleic acid sequence data having the same taxonomic name and / or taxonomic ID to satisfy at least one of the following predetermined alignment criteria; and (c-2) selecting the nucleic acid sequence of the highest-ranking nucleic acid sequence data among the aligned nucleic acid sequence data as the taxonomic representative sequence, and the predetermined alignment criteria include the following: (i) alignment according to the assembly level of the nucleic acid sequence data; the assembly level is ranked in the order of complete genome, chromosome, scaffold, and contig; and (ii) alignment according to whether the nucleic acid sequence data is included in a RefSeq (Reference Sequence) database; (iii) sorting based on whether the name of the nucleic acid molecule listed in the descriptor of the nucleic acid record containing the nucleic acid sequence data matches at least one of the name of the received target nucleic acid molecule and the collected synonyms; the case of a match is higher than the case of a non-match; (iv) sorting based on the length of the nucleic acid sequence data; the longer the length, the higher the rank; (v) sorting based on whether a host is listed in the descriptor of the nucleic acid record containing the nucleic acid sequence data; the case where the host of interest for the organism of interest is listed is higher than the case where it is not listed, and the case where it is not listed is higher than the case where an organism different from the host of interest is listed in the host; and (vi) sorting based on the registration date or modification date of the nucleic acid record containing the nucleic acid sequence data.The higher the date of registration or modification, the higher the rank; (vii) sorted according to the alphabet of the access number of the nucleic acid sequence data; the earlier the alphabet of the access number, the higher the rank. Claim 11 A method according to claim 1, wherein the selection of the group representative sequence in step (d) is carried out by a method comprising the following steps: (d-1) aligning the selected taxonomic representative sequence to satisfy at least one of the following predetermined alignment criteria; (d-2) selecting the highest taxonomic representative sequence among the aligned taxonomic representative sequences; and (d-3) grouping taxonomic representative sequences having homology greater than or equal to a predetermined value with respect to the highest taxonomic representative sequence and selecting the highest taxonomic representative sequence in each group as the group representative sequence, and the alignment criteria include the following: (i) alignment according to the assembly level of the selected taxonomic representative sequence; the assembly level is ranked in the order of complete genome, chromosome, scaffold, and contig; and (ii) the number of nucleic acid sequence data having the same taxonomic name and / or taxonomic identification symbol as the selected taxonomic representative sequence; The higher the number, the higher the rank; (iii) sorting according to whether a host is listed in the descriptor of the nucleic acid record containing the selected taxonomic representative sequence; the case where the host of interest for the organism of interest is listed in the host is higher in rank than the case where it is not listed, and the case where it is not listed is higher in rank than the case where an organism other than the host of interest is listed in the host; (iv) sorting according to the alphabet of the accession number of the selected taxonomic representative sequence; the earlier the alphabet of the accession number, the higher the rank. Claim 12 A method according to claim 1, wherein the method further comprises the step of (f) providing the nucleic acid sequence data relating to the received organism of interest among the nucleic acid sequence data sets provided in step (e) as a target nucleic acid sequence data set relating to a target nucleic acid molecule. Claim 13 A method according to claim 12, wherein the target nucleic acid sequence data set provided in step (f) has homology of at least a predetermined value with respect to the target nucleic acid sequence data set with respect to at least one representative sequence among a group representative sequence and a taxonomic representative sequence. Claim 14 The method of claim 1, wherein the method further comprises the step of (g) providing nucleic acid sequence data that is not related to the received organism of interest among the nucleic acid sequence data sets provided in step (e) as a non-target nucleic acid sequence data set for non-target nucleic acid molecules. Claim 15 A method according to claim 14, wherein the non-target nucleic acid sequence data set provided in step (g) satisfies at least one of the following homology criteria: (i) the non-target nucleic acid sequence data set has homology greater than a predetermined value with respect to a portion of the sequence region of at least one representative sequence among the group representative sequence and the taxonomic representative sequence; (ii) the non-target nucleic acid sequence data set has homology greater than a predetermined value with respect to at least one representative sequence among the group representative sequence and the taxonomic representative sequence; and (iii) the non-target nucleic acid sequence data set having the homology criterion of (i) has the homology criterion of (ii). Claim 16 In claim 12, the method is characterized by additionally comprising the following steps: (h) collecting nucleic acid sequence data having homology greater than or equal to a predetermined value with respect to the group representative sequence and providing a nucleic acid sequence data set; and (j) classifying the target nucleic acid sequence data of the group representative sequence and the target nucleic acid sequence data belonging to the same group as the group representative sequence in the target nucleic acid sequence data set of step (f) as design-excluded target nucleic acid sequence data when the taxonomic name and / or taxonomic identification symbol of the organism for the group representative sequence satisfies one of the following predetermined criteria; and, the predetermined criteria include the following: (i) when there is no nucleic acid sequence data corresponding to the taxonomic name and / or taxonomic identification symbol of the organism for the group representative sequence in the nucleic acid sequence data set provided in step (h), and only nucleic acid sequence data corresponding to the taxonomic name and / or taxonomic identification symbol of an organism different from the organism for the group representative sequence exists; (ii) if the homology of the nucleic acid sequence data corresponding to the taxonomic name and / or taxonomic identifier of the organism for the group representative sequence in the nucleic acid sequence data set provided in step (h) is lower compared to the homology of the nucleic acid sequence data corresponding to the taxonomic name and / or taxonomic identifier of the organism for the group representative sequence and another organism; (iii) if the target nucleic acid sequence data corresponding to the taxonomic name and / or taxonomic identifier of the organism for the group representative sequence does not exist in the nucleic acid sequence data set provided in step (h), and the ratio of nucleic acid sequence data corresponding to an organism corresponding to a taxonomic name and / or taxonomic identifier higher than the organism for the group representative sequence or an organism corresponding to a taxonomic name and / or taxonomic identifier lower than that is less than a predetermined value for the nucleic acid sequence data set provided in step (h);and (iv) where the nucleic acid sequence data set provided in step (h) is a nucleic acid sequence data set of an organism for the group representative sequence, but the name of the target nucleic acid molecule is not listed in the descriptors of the nucleic acid records containing the nucleic acid sequence data set, or is different from at least one of the name of the target nucleic acid molecule and the collected synonyms.; Claim 17 A method according to claim 1, wherein the method further comprises the following steps after step (e): (e-1) sorting the provided design nucleic acid sequence data set according to a biosample identifier to select nucleic acid sequence data having the same biosample identifier; (e-2) sorting the selected nucleic acid sequence data to satisfy at least one of the following alignment criteria; (e-3) selecting the nucleic acid sequence of the top nucleic acid sequence data among the sorted nucleic acid sequence data; and (e-4) removing nucleic acid sequence data other than the top nucleic acid sequence data from the design nucleic acid sequence data set, and the alignment criteria include: (i) sorting the nucleic acid sequences included in the provided design nucleic acid sequence data set according to the assembly level; The assembly levels are ranked in the order of complete genome, chromosome, scaffold, and contig, with the highest ranking being (ii) sorted according to whether the nucleic acid sequences included in the provided design nucleic acid sequence data set are included in the RefSeq (Reference Sequence) database; the ranking is higher when the nucleic acid sequence data is included in the RefSeq database than when it is not included. Claim 18 In claim 14, the method is characterized by additionally comprising the following steps: (k) collecting nucleic acid sequence data having homology greater than or equal to a predetermined value with respect to the non-target nucleic acid sequence of an organism in the non-target nucleic acid sequence data set and providing the nucleic acid sequence data set; and (l) classifying the non-target nucleic acid sequence data of the organism in the non-target nucleic acid sequence data set of step (k) as design-excluded non-target nucleic acid sequence data when the taxonomic name and / or taxonomic identification symbol of the organism for the non-target nucleic acid sequence satisfies one of the following predetermined criteria; and, the predetermined criteria include the following: (i) when there is no nucleic acid sequence data corresponding to the taxonomic name and / or taxonomic identification symbol of the organism for the non-target nucleic acid sequence in the nucleic acid sequence data set provided in step (k), and only nucleic acid sequence data corresponding to the taxonomic name and / or taxonomic identification symbol of an organism other than the organism exists; (ii) if the homology of the non-target nucleic acid sequence data corresponding to the taxonomic name and / or taxonomic identification symbol of the organism for the non-target nucleic acid sequence in the nucleic acid sequence data set provided in step (k) is lower compared to the homology of the nucleic acid sequence data corresponding to the taxonomic name and / or taxonomic identification symbol of the organism and other organisms; and (iii) if there is no non-target nucleic acid sequence data corresponding to the taxonomic name and / or taxonomic identification symbol of the organism for the non-target nucleic acid sequence in the nucleic acid sequence data set provided in step (k), and the ratio of nucleic acid sequence data corresponding to an organism corresponding to a taxonomic name and / or taxonomic identification symbol higher than the organism or an organism corresponding to a taxonomic name and / or taxonomic identification symbol lower than that organism is less than a predetermined value for the nucleic acid sequence data set provided in step (k). Claim 19 A computer-readable recording medium storing instructions for executing on a processor a method for providing a nucleic acid sequence data set for designing an oligonucleotide used to detect a target nucleic acid molecule of an organism of interest, wherein the method comprises the following steps: (a) receiving the name of the target nucleic acid molecule and the name of the organism of interest, and retrieving synonyms for the target nucleic acid molecule of the organism of interest; (b) collecting nucleic acid sequence data contained in nucleic acid records; Each of the above nucleic acid records relates to the organism of interest and includes a descriptor in which the name of the target nucleic acid molecule and at least one of the collected synonyms are listed; (c) sorting the collected nucleic acid sequence data according to a taxonomic name and / or a taxonomic ID and selecting a taxonomic representative sequence from among the nucleic acid sequence data having the same taxonomic name and / or taxonomic ID; (d) grouping the selected taxonomic representative sequences according to homology and selecting a group representative sequence from each group; and (e) collecting nucleic acid sequence data having homology of a predetermined value or more with the group representative sequence and providing it as a nucleic acid sequence data set for designing the oligonucleotide.
Citation Information
Patent Citations
Method for providing target nucleic acid sequence data set of target nucleic acid molecule
WO2019212238A1