Filtering method of nucleic acid sequence information

The method addresses the challenge of removing host sequence information from nucleic acid sequence information by using a closely related organism's genome sequence as a reference for mapping, enhancing the efficiency of biodiversity analysis even when the host organism's genome is not fully sequenced.

JP2025079493APending Publication Date: 2025-05-22KK TOYOTA CHUO KENKYUSHO
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023192201
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-10
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

Existing methods struggle to efficiently remove host sequence information from nucleic acid sequence information when analyzing biological samples from environments where the entire genome sequence of the host organism has not been decoded.

Method used

A method involving the extraction of genomic DNA from a specific organism, generation of genomic reads, de novo assembly to create contigs, homology search to identify closely related organisms, and mapping of nucleic acid reads using the genome sequence of a closely related organism as a reference to filter out host sequence information.

Benefits of technology

This method allows for efficient removal of host sequence information, even when the entire genome sequence of the host organism is unknown, thereby improving the efficiency of analyzing nucleic acid sequence information related to microorganisms and viruses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025079493000001_ABST
    Figure 2025079493000001_ABST
Patent Text Reader

Abstract

To efficiently remove a specific organism-derived nucleic acid sequence from nucleic acid sequence information to be analyzed in the case where a specific organism whose entire genome sequencing has not been completed is used an organism-derived nucleic acid source.SOLUTION: In a filtering method of nucleic acid sequence information, a genome DNA is extracted from a first specific organism whose entire genome sequencing has not been completed, a genome read is acquired by using the extracted genome DNA, a contig is created by using the genome read, homology search by using database is performed to specify a closely related organism which belongs to the same family as the first specific organism, organism-derived nucleic acid attached to specific organisms such as the first specific organism is recovered from the specific organisms, a nucleic acid read is acquired from the recovered organism-derived nucleic acid, mapping is performed by using a genome sequence of the closely related organism as a reference sequence, and a residual nucleic acid read other than the nucleic acid read aligned on the reference sequence is acquired as nucleic acid sequence information to be analyzed.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to methods for filtering nucleic acid sequence information. [Background technology]

[0002] In order to investigate the diversity of organisms living in a specific environment, environmental nucleic acids including biological nucleic acids are collected from the environment to be investigated, and the collected biological nucleic acids are analyzed. In order to analyze biological nucleic acids, it is necessary to collect biological nucleic acids from a relatively wide range that constitutes the specific environment. However, when attempting to collect biological nucleic acids from environmental water such as pond water or the atmosphere, the biological nucleic acids are present in a dilute state, which may make it difficult to sufficiently collect and analyze the biological nucleic acids. As one method for solving such a problem, for example, a method of collecting biological nucleic acids that are concentrated by being attached to or taken in by organisms that move in environments such as environmental water and the atmosphere, such as fish and birds, or substances derived from such organisms, and analyzing the collected biological nucleic acids is considered. However, biological nucleic acids collected from organisms (hereinafter also referred to as "host organisms") that are the source of the biological nucleic acids as described above will contain a large amount of nucleic acids derived from the host organism. In order to investigate biodiversity, it is important to analyze the diversity of, for example, microorganisms and viruses. Therefore, in order to reduce analytical processing and improve analytical efficiency, it is desirable to analyze the nucleic acid sequence information of the recovered biological nucleic acid by excluding the sequence information of the nucleic acid derived from the host organism (hereinafter also referred to as "host sequence information").

[0003] When collecting and analyzing biological nucleic acids from a specific organism, a method for removing the nucleic acid sequence information of the specific organism from which the biological nucleic acids are collected is known, for example, a technique for analyzing the microbiome (microbial population) of a bee. Specifically, a technique is known in which DNA is extracted from the surface or inside the body of a bee, the base sequence is read, the obtained base sequence is mapped to a known bee genome sequence, and the mapped base sequence is removed, thereby removing the nucleic acid sequence information of the bee (the nucleic acid sequence information of the specific organism) from the nucleic acid sequence information of the extracted DNA, and the microbiome is analyzed using the remaining nucleic acid sequence information (for example, see Non-Patent Document 1).

[0004] A similar technique is known in which nucleic acid sequence information is obtained by recovering nucleic acid of biological origin from humans, known human genome sequence information is removed from the obtained nucleic acid sequence information, and the remaining nucleic acid sequence information is used to analyze nucleic acid sequence information related to pathogens, enterobacteria, etc. in the nucleic acid of biological origin recovered from humans. For example, a technique has been proposed in which a nucleic acid sample is recovered from a patient to obtain nucleic acid sequence information, known human genome sequence information is removed from the obtained nucleic acid sequence information, and the remaining nucleic acid sequence information is used to analyze the novel coronavirus (SARS-CoV-2) (see, for example, Non-Patent Document 2). Furthermore, as a related technique, a tool (framework) for removing sequence information of human-derived DNA, which is the source of nucleic acid recovery, from nucleic acid sequence information has been proposed (see, for example, Non-Patent Document 3). [Prior art documents] [Non-patent literature]

[0005] [Non-Patent Document 1] Katherine D. Chau et al., "Integrative population genetics and metagenomics reveals urbanization increases pathogen loads and decreases connectivity in a wild bee", Global Change Biology, (2023) 29, 4193-4211. DOI:10.1111 / gcb.16757 [Non-Patent Document 2] Sun Zhaoyang et al., "Clinical characteristics of the host DNA-removed metagenomic next-generation sequencing technology for tectiong SARS-CoV-2, revealing host local immune signaling and assisting genomic epidemiology", Frontiers in Immunology, 15 November 2022, DOI:10.3389 / fimmu.2022.1016440 [Non-Patent Document 3] Robert Schmieder et al., "Fast Identification and Removal of Sequence Contamination from Genomic and Metagenomic Datasets", PLoS ONE, March 2011;6(3). [Summary of the Invention] [Problems to be Solved by the Invention]

[0006] In the above-mentioned conventionally known techniques, the entire genome sequence of a specific organism (bee or human) that is the source of biological nucleic acid is known, and it is relatively easy to remove the genome sequence information of such a specific organism from the nucleic acid sequence information to be analyzed. In contrast, when investigating the biodiversity of an arbitrary environment, it is assumed that the entire genome sequence of a host organism selected from organisms living in the environment has not been decoded. In such a case, it is difficult to sufficiently remove the host sequence information from the nucleic acid sequence information to be analyzed by applying the conventionally known techniques. If the host sequence information cannot be removed from the nucleic acid sequence information to be analyzed because the entire genome sequence of the host organism has not been decoded, it becomes necessary to analyze the entire nucleic acid sequence information including the host sequence information. However, for example, the sequence information of nucleic acid recovered from the body surface of a fish may contain about 80% sequences derived from the host organism, which reduces the efficiency of analysis. Therefore, even when a specific organism whose entire genome sequence has not been decoded is used as the source of biological nucleic acid, a technique for efficiently removing (filtering) the nucleic acid sequence of the specific organism from the nucleic acid sequence information to be analyzed has been desired. [Means for solving the problem]

[0007] The present disclosure can be realized in the following forms. (1) According to one embodiment of the present disclosure, a method for filtering nucleic acid sequence information is provided. The method for filtering nucleic acid sequence information includes extracting genomic DNA from a first specific organism whose entire genome sequence has not been decoded, obtaining genomic reads, which are nucleic acid sequences obtained by fragmenting the genomic DNA using the extracted genomic DNA, and performing decode using the genomic reads. A novo assembly is performed to generate contigs, which are partial sequences of the genomic DNA of the first specific organism, and a homology search is performed using a database containing genome sequence information and sequence information of the contigs to identify closely related organisms, including an organism that belongs to the same family as the first specific organism and has the largest number of contigs that satisfy a predetermined criterion for determining homology, among the organisms for which the genome sequence information is stored in the database; biological nucleic acids attached to the first specific organism or a specific organism that is a second specific organism that belongs to the same family as the first specific organism and whose entire genome sequence has not been deciphered, or biological nucleic acids present in the body of the specific organism are recovered from the specific organism; nucleic acid reads whose base sequences have been read by next-generation sequencing (NGS) are obtained from the recovered biological nucleic acids; mapping is performed using the nucleic acid reads with the genome sequence of the closely related organism as a reference sequence; and the remaining nucleic acid reads other than the nucleic acid reads aligned on the reference sequence are obtained as nucleic acid sequence information to be analyzed. According to this embodiment of the method for filtering nucleic acid sequence information, when analyzing nucleic acid reads obtained from biological nucleic acids recovered from specific organisms whose entire genome sequence has not been deciphered, the genome sequence of a closely related organism that belongs to the same family as the first specific organism and is considered to show high homology with the genome sequence of the first specific organism is used as a reference sequence for mapping, and unaligned nucleic acid reads are obtained, thereby filtering the nucleic acid reads. Therefore, even if the entire genome sequence of the specific organisms that are the source of biological nucleic acids has not been deciphered, when analyzing biological nucleic acids recovered from the specific organisms, the genome sequence information of the specific organisms that are the source of biological nucleic acids can be efficiently removed. Therefore, analysis can be efficiently performed by excluding nucleic acid sequence information derived from the biological nucleic acid source (specific organisms) that occupies a relatively large proportion of biological nucleic acids. (2) In the method for filtering nucleic acid sequence information of the above aspect, the homology search may be performed using contigs among the generated contigs, the average coverage of the genome reads of which is equal to or greater than a predetermined reference value. With this configuration, the homology search can be performed using contigs that are more accurate as constituting a part of the true genome sequence of the first specific organism. (3) In the method for filtering nucleic acid sequence information of the above embodiment, fish may be used as the specific organism. With this configuration, it is possible to efficiently remove genome sequences derived from the specific organism, which is the source of biological nucleic acid, from the sequence information of the nucleic acid read to be analyzed. (4) In the above-described method for filtering nucleic acid sequence information, the de novo assembly may be performed using genome reads in which the total amount of nucleic acid sequence information obtained by sequencing in shotgun metagenomic analysis is 2 Gb or more. With this configuration, it is possible to ensure a sufficient amount of sequence information of the genome of the first specific organism to be subjected to the de novo assembly. (5) In the method for filtering nucleic acid sequence information of the above embodiment, three or more organisms may be identified as the closely related organisms in order of the number of contigs that satisfy the predetermined criteria, and the genome sequences of the identified three or more closely related organisms may be used as the reference sequences. With this configuration, it is possible to increase the efficiency of removing genome sequence information of a specific organism that is a source of organism-derived nucleic acid from the nucleic acid reads obtained from the organism-derived nucleic acid. (6) In the above-described method for filtering nucleic acid sequence information, the specific organism may be a mobile organism, and the nucleic acid lead may be obtained from a biological nucleic acid attached to the specific organism, or from a biological nucleic acid contained in a substance discharged from the specific organism. In this configuration, biological nucleic acids in the environment in which the specific organism lives are concentrated by attaching to the specific organism, and the environment in which the specific organism lives can be efficiently analyzed by using the nucleic acid lead obtained by obtaining biological nucleic acids from the specific organism. The present disclosure can be realized in various forms other than those described above, for example, in the form of a method for analyzing nucleic acid derived from an organism, a method for investigating biodiversity, and the like. [Brief description of the drawings]

[0008] [Figure 1] 4 is a flowchart showing a method for filtering nucleic acid sequence information according to the first embodiment. [Diagram 2] 10 is a flowchart showing a method for filtering nucleic acid sequence information according to a second embodiment. [Diagram 3] An explanatory diagram showing the BLAST search results for contigs that hit bony fish. [Figure 4] An explanatory diagram showing the results of mapping RNA-seq reads on the body surface of the bitterling. [Diagram 5] An explanatory diagram showing the results of mapping RNA-seq reads on the body surface of the bitterling. [Figure 6] An explanatory diagram showing BLAST search results for divided genome read groups. [Figure 7] An explanatory diagram showing BLAST search results for divided genome read groups. [Figure 8] An explanatory diagram showing BLAST search results for divided genome read groups. [Figure 9] An explanatory diagram showing BLAST search results for divided genome read groups. [Figure 10] An explanatory diagram showing the results of mapping various RNA-seq reads. [Figure 11] An explanatory diagram showing the phylogenetic tree of 39 major species of freshwater fish. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0009] A. First embodiment: 1 is a flowchart showing a method for filtering nucleic acid sequence information according to a first embodiment of the present disclosure. The method for filtering nucleic acid sequence information according to the first embodiment is a method for obtaining a filtering effect similar to that obtained by removing genome sequence information of a first specific organism from sequence information of the collected nucleic acid when collecting nucleic acid derived from an organism from a first specific organism whose entire genome sequence (sequences of chromosomal DNA and mitochondrial DNA, and in the case of a plant, further including a sequence of chloroplast DNA) has not yet been deciphered and analyzing the collected nucleic acid derived from the organism.

[0010] The method for filtering nucleic acid sequence information of the present embodiment can be used, for example, to collect environmental nucleic acids including biological nucleic acids from an environment to be investigated, and to analyze the collected environmental nucleic acids, thereby investigating the diversity of organisms living in the environment to be investigated. Specifically, the method for filtering nucleic acid sequence information of the present embodiment can be applied when using a first specific organism living in the environment to be investigated as a source of biological nucleic acids (host organism) and collecting and analyzing biological nucleic acids from the host organism. As the host organism, for example, a mobile organism that moves in an environment such as environmental water or air, such as fish, birds, or insects, can be used. The operation of recovering environmental nucleic acids from a host organism can be performed, for example, by recovering environmental nucleic acids attached to the host organism from the body surface of the host organism or from substances detached from the body surface (such as fish scales or bird feathers). Alternatively, the operation of recovering environmental nucleic acids from a host organism can be performed by recovering environmental nucleic acids taken into the host organism from internal substances such as the digestive tract contents of the host organism, or from excrement of the host organism. Furthermore, when a non-mobile plant or the like is used as the host organism, biological nucleic acid may be recovered from seeds or spores that are released from the host organism and move around in the environment, or from parts of the organism (such as fallen leaves) that detach from the host organism and move around in the environment.

[0011] Furthermore, the method for filtering nucleic acid sequence information of this embodiment can be used to investigate the state of a first specific organism to be investigated by recovering biological nucleic acid from the first specific organism to be investigated and analyzing the recovered biological nucleic acid. Specifically, for example, the state of the first specific organism can be investigated by recovering nucleic acid from the body surface or inside the body of the first specific organism to be investigated and analyzing the microflora on the body surface or inside the body of the first specific organism. Alternatively, the method for filtering nucleic acid sequence information of this embodiment can be used to determine the presence or absence of infection of a specific pathogenic microorganism in the first specific organism by detecting the nucleic acid sequence of the specific pathogenic microorganism in the biological nucleic acid recovered from the first specific organism.

[0012] By applying the nucleic acid sequence information filtering method of this embodiment to the above-mentioned surveys, etc., when recovering biological nucleic acid from a first specific organism (host organism) whose entire genome sequence has not been deciphered, it becomes possible to exclude the nucleic acid sequence information of the host organism (host sequence information) from the nucleic acid sequence information of the recovered biological nucleic acid.

[0013] When performing the method for filtering nucleic acid sequence information shown in Fig. 1, first, genomic DNA is extracted from a first specific organism whose entire genome sequence has not been deciphered (step T100). Here, "whole genome sequence has not been deciphered" for the first specific organism means that no or only a small amount of the genome sequence of the first specific organism has been deciphered, and there is essentially no database that can be used to obtain genome sequence information of the first specific organism. Extraction of genomic DNA in step T100 can be performed from a part of the body tissue of the first specific organism using a well-known DNA extraction method.

[0014] After the genomic DNA of the first specific organism is extracted, the extracted genomic DNA is fragmented to obtain a nucleic acid sequence, which is a read (hereinafter also referred to as a "genomic read") (step T110). The genomic read can be obtained by a shotgun metagenomic analysis in which the extracted genomic DNA is fragmented to a size suitable for sequencing, and the obtained fragments are randomly sequenced.

[0015] Thereafter, de novo assembly is performed using the genome reads obtained in step T110, and a contig, which is a partial sequence of the genome DNA of the first specific organism, is generated by connecting multiple genome reads (step T120). De novo assembly is a well-known sequence generation method that assembles sequence reads to construct a new sequence without using a reference sequence. From the viewpoint of ensuring a sufficient amount of sequence information of the genome of the first specific organism to be subjected to de novo assembly, it is preferable to perform de novo assembly using genome reads in which the total amount of nucleic acid sequence information obtained by sequencing in the shotgun metagenomic analysis is 2 Gb or more, more preferably 5 Gb or more of genome reads are used, and even more preferably 10 Gb or more of genome reads are used.

[0016] When generating contigs from genome reads in step T120, the genome read data to be used may be pre-processed. For example, since adapter sequences are generally added to the genome reads obtained in step T110, the adapter sequences may be removed from the genome reads as pre-processing prior to generating contigs.

[0017] In addition, a process for selecting genome reads to be used may be performed as a pre-processing performed on genome read data. As a method for selecting genome reads, for example, genome reads with a relatively low quality score (Q value) among the genome reads obtained in step T110 may be removed (quality filtering). The quality score indicates the probability of an error occurring in reading the base sequence, and for example, a quality score of 20 (Q20) indicates that the probability of an error occurring is 1 in 100, and a quality score of 30 (Q30) indicates that the probability of an error occurring is 1 in 1000 and the reliability is 99.9%. Therefore, for example, by removing at least some genome reads with a quality score of less than 20 to generate contigs, the accuracy that the obtained contigs constitute a part of the true genome sequence of the first specific organism can be improved.

[0018] In addition, as a process for selecting genome reads in step T120, genome reads shorter than a predetermined standard length may be removed from genome reads used to generate contigs. This is because the shorter the genome read length, the higher the possibility of incorrect alignment occurring when generating contigs. In order to ensure that the sequence of the contig generated in step T120 is a part of the true genome sequence of the first specific organism, genome reads of, for example, 30 bp or less may be removed, genome reads of 50 bp or less may be removed, or genome reads of 70 bp or less may be removed.

[0019] After step T120, a homology search is performed using a database containing genome sequence information and sequence information of the contigs generated in step T120 to identify closely related organisms that belong to the same family as the first specific organism (step T130). The homology search performed in step T130 can be easily performed by, for example, a blast search. The closely related organisms identified as a result of the homology search in step T130 refer to organisms that belong to the same family as the first specific organism among organisms whose genome sequence information is stored in the database, and that include the organism that contains the organism with the largest number of contigs that satisfy a predetermined criterion for determining that there is homology. As the above-mentioned "predetermined criterion for determining that there is homology," an E-value can be used when a blast search is performed as the homology search, and it may be determined that the above criterion is satisfied when, for example, the E-value is 0.0001 or less. In a homology search, a specific contig that satisfies the above-mentioned predetermined criterion is also referred to as "hitting." An organism that has a genome sequence that has the largest number of contigs that hit in the homology search is also referred to as the "closest related organism." A large number of contig hits in a homology search is considered to reflect a high degree of homology between the genome sequence of an organism with a large number of contig hits and the genome sequence of the first specified organism.

[0020] In step T130, it is sufficient to identify one or more organisms including the most related organism as the closely related organism, preferably to identify two or more organisms including the most related organism, and more preferably to identify three or more organisms including the most related organism. When multiple organisms are identified as closely related organisms, the closely related organisms may be selected in descending order of the number of contigs that satisfy (hit) the above-mentioned predetermined criteria. In addition, it is preferable that the closely related organisms identified in step T130 are organisms whose entire genome sequence has been deciphered and whose sequence information of the entire genome sequence is stored in a database.

[0021] In the above-mentioned step T120, a large number of contigs are generated according to the amount of sequence information of the genome reads used, and therefore, in step T130, prior to the homology search, contigs to be used in the homology search may be narrowed down (selected) from the contigs generated in step T120. For example, in order to perform a homology search using a contig that is considered to have a higher accuracy of constituting a part of the true genome sequence of the first specific organism, it is desirable to select contigs whose average coverage by genome reads is equal to or greater than a predetermined reference value from among the contigs generated in step T120 and use them in the homology search. Specifically, for example, it is preferable to use contigs whose average coverage by genome reads is 500 or more, more preferably 1000 or more, and even more preferably 2000 or more. Here, the average coverage refers to the average number of genome reads overlapped on this contig to obtain one contig. Specifically, the average coverage in each contig is calculated by calculating the coverage (number of overlapping genomic reads) for every position (each base) in each contig, and then calculating the average coverage for each base for each contig.

[0022] Furthermore, when performing the homology search in step T130, in addition to or instead of the above-mentioned selection based on the average coverage, the contigs used in the homology search may be selected by other methods. For example, the contigs may be selected based on the length of the contigs, and specifically, only contigs longer than a predetermined lower limit may be used in the homology search in step T130. The shorter the length of the contigs, the higher the possibility that they will be determined to have homology at a position other than the truly corresponding position in the genome sequence to be searched, so the accuracy of the homology search can be improved by removing relatively short contigs.

[0023] In addition, in the method for filtering nucleic acid sequence information of this embodiment, apart from the operations from the extraction of genomic DNA from the first specific organism in step T100 to the identification of closely related organisms in step T130, biological nucleic acids are collected from the first specific organism, and nucleic acid reads are obtained from the collected biological nucleic acids (step T140). Step T140 is a step for obtaining biological nucleic acids to be analyzed. The operation of collecting biological nucleic acids from the first specific organism in step T140 can be performed by collecting biological nucleic acids attached to the first specific organism, or biological nucleic acids present in the body of the first specific organism. Specifically, biological nucleic acids can be collected from the body surface of the first specific organism, substances detached from the body surface, the digestive tract contents of the first specific organism, excreta of the first specific organism, etc. In step T140, after the biological nucleic acids are collected, nucleic acid reads whose base sequences are read by next-generation sequencing (NGS) are obtained from the collected biological nucleic acids.

[0024] The biological nucleic acid recovered in step T140 may be deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). A nucleic acid read obtained by recovering DNA as biological nucleic acid and comprehensively determining the base sequence of the DNA by next-generation sequencing is also called a DNA-seq read, and a nucleic acid read obtained by recovering RNA as biological nucleic acid and comprehensively determining the base sequence of the RNA by next-generation sequencing is also called an RNA-seq read. When DNA is used as biological nucleic acid, DNA is more stable than RNA, so by analyzing biological nucleic acid, it is possible to detect organisms with a lower abundance ratio, and it is also possible to obtain analysis results that reflect the relatively long-term state in the environment from which the biological nucleic acid was obtained. When RNA is used as biological nucleic acid, RNA is more unstable than DNA, so by analyzing biological nucleic acid, it is possible to obtain analysis results that reflect the state at the time of collection of biological nucleic acid in the environment from which the biological nucleic acid was obtained, and it is also possible to analyze the state of genes that are actually expressed.

[0025] After acquiring the nucleic acid reads in step T140, the genome sequence of the closely related organism identified in step T130 is used as a reference sequence to perform mapping using the nucleic acid reads acquired in step T140. Then, the remaining nucleic acid reads other than the nucleic acid reads aligned on the reference sequence are acquired as the nucleic acid sequence information to be analyzed, thereby completing the filtering of the nucleic acid sequence information (step T150). The operation of aligning the nucleic acid reads on the reference sequence can be performed using a computer program by any known method. The above mapping is performed based on the sequence homology between the reference sequence and the nucleic acid read. In step T130, when multiple organisms are identified as closely related organisms, the genome sequences of all organisms identified as closely related organisms are used as reference sequences, and all nucleic acid reads aligned on any of the reference sequences are removed from the analysis target.

[0026] According to the method for filtering nucleic acid sequence information of the first embodiment configured as described above, when analyzing nucleic acid reads obtained from biological nucleic acid recovered from a first specific organism whose entire genome sequence has not been decoded, the genome sequence of a closely related organism that belongs to the same family as the first specific organism and is considered to show high homology with the genome sequence of the first specific organism is used as a reference sequence for mapping, and unaligned nucleic acid reads are obtained, thereby filtering the nucleic acid reads. Therefore, even if the entire genome sequence of the first specific organism, which is a host organism, is not decoded, when analyzing biological nucleic acid recovered from the first specific organism, the genome sequence information of the host organism can be efficiently removed from the sequence information of the recovered biological nucleic acid. As a result, analysis can be efficiently performed by excluding the nucleic acid sequence information of the host organism, which occupies a relatively large proportion of the biological nucleic acid recovered from the host organism, and analysis of nucleic acid derived from organisms other than the host organism (for example, analysis of the diversity of microorganisms, viruses, etc.) can be easily performed.

[0027] When applying the method for filtering nucleic acid sequence information of this embodiment, if a fish is used as the first specific organism, the host organism, the operation of removing the genome sequence derived from the host organism from the nucleic acid sequence information to be analyzed can be performed more efficiently. For example, the sequence information of the nucleic acid recovered from the body surface of the fish may contain about 80% of the sequence derived from the host organism. When setting the closely related organisms described above having a genome sequence that is considered to show high homology with the genome sequence of the host organism, if appropriate settings are made, for example, by setting three or more closely related organisms, and filtering of the organism-derived nucleic acid derived from the host organism is performed, it becomes possible to remove, for example, 80% or more of the sequence from the sequence information of the organism-derived nucleic acid recovered from the host organism. In other words, it is considered possible to perform filtering that removes most of the genome sequences of the fish, the host organism.

[0028] B. Second embodiment: Fig. 2 is a flow chart showing a method for filtering nucleic acid sequence information according to the second embodiment. In Fig. 2, steps common to the first embodiment are given the same step numbers. The method for filtering nucleic acid sequence information according to the second embodiment can be used for the same applications and purposes as the method for filtering nucleic acid sequence information according to the first embodiment. In the method for filtering nucleic acid sequence information according to the second embodiment, steps T100 to T130 are common to the first embodiment, but steps T240 and T250 are performed instead of steps T140 and T150.

[0029] The method for filtering nucleic acid sequence information of the second embodiment can be used for the same purpose as the method for filtering nucleic acid sequence information of the first embodiment, but a second specific organism different from the first specific organism is used to obtain nucleic acid sequence information to be analyzed. That is, in the method for filtering nucleic acid sequence information of the first embodiment, the organism from which genomic DNA is extracted in step T100 to identify a closely-related organism to be used for filtering, and the host organism from which biological nucleic acid is recovered in step T140 to obtain a nucleic acid lead to be analyzed, are both the same first specific organism. In contrast, in the second embodiment, biological nucleic acid is recovered using a second specific organism different from the first specific organism in step T240, and the nucleic acid lead to be analyzed is obtained. This second specific organism is an organism that belongs to the same family as the first specific organism from which genomic DNA is extracted in step T100, and whose entire genome sequence has not been deciphered. The first specific organism used as a biological nucleic acid source (host organism) in step T140 of the first embodiment and the second specific organism used as a biological nucleic acid source in step T240 of the second embodiment are collectively referred to as "specific organisms".

[0030] In the second embodiment, the operation of obtaining nucleic acid reads from the second specified organism in step T240 can be performed in the same manner as in step T140. Then, in step T250, using genomic DNA extracted from the first specified organism, and using the genomic sequence of the closely related organism specified in step T130 as a reference sequence, the nucleic acid reads derived from the second specified organism obtained in step T240 are mapped, and the sequence information of the organism-derived nucleic acid recovered from the second specified organism is filtered.

[0031] With this configuration, when recovering and analyzing biological nucleic acids using a second specific organism whose entire genome sequence has not been decoded as a host organism, it is possible to filter biological nucleic acids recovered from the second specific organism without performing the operation of extracting the genome DNA of the host organism itself (second specific organism) and setting a closely related organism thought to show high homology to the genome sequence of the second specific organism as in the first embodiment. That is, for a first specific organism whose entire genome sequence has not been decoded and belongs to the same family as the second specific organism, if a closely related organism of the first specific organism is identified in step T130, mapping is performed using the genome sequence of this closely related organism as a reference sequence to remove the aligned nucleic acid reads, thereby making it possible to efficiently remove the genome sequence of the second specific organism, which is the host organism, from the biological nucleic acids recovered from the second specific organism. Therefore, for example, if one type of organism is used as a first specified organism and closely related organisms are identified by carrying out steps T100 to T130, then when analyzing organism-derived nucleic acids using another organism that belongs to the same family as the first specified organism and whose entire genome sequence has not yet been deciphered as a second specified organism that is the source of organism-derived nucleic acids, it becomes possible to filter the nucleic acid sequence information by omitting the processing of steps T100 to T130. EXAMPLES

[0032] [Example corresponding to the first embodiment] As the first specified organism, we used the bitterling, whose entire genome sequence has not yet been deciphered, and implemented and examined the method for filtering nucleic acid sequence information shown in Figure 1.

[0033] <Extraction of genomic DNA from the Japanese bitterling: Step T100> The tail fin of the first specified organism, the bitterling, was cut into approximately 5 mm squares and suspended in 1 mL of DNA / RNA Shield solution, a preservation reagent for DNA and RNA samples. Proteinase K (Takara Bio Inc.), a proteolytic enzyme, was added to the resulting suspension in an amount of 1 / 40 of the suspension, and the suspension was incubated at 55°C for 15 minutes to decompose mucin, which is abundant in the mucus, in order to improve the efficiency of DNA extraction. DNA was then extracted using the ZymoBIOMICS DNA Miniprep Kit (Zymo Research) according to the protocol. The DNA concentration in the extract was measured by fluorometric quantification using the QuantiFluor ONE dsDNA System (Promega).

[0034] <Shotgun metagenomic analysis of the Japanese bitterling and identification of candidates for closely related organisms: Steps T110 to T130> A shotgun metagenomic analysis was performed using DNA extracted from the Japanese bitterling. Shotgun metagenomic sequencing by next-generation sequencing (NGS) was outsourced (requested from Seibu Giken Co., Ltd.). In the shotgun metagenomic analysis, genome reads with a total amount of nucleic acid sequence information of 70 Gb were obtained.

[0035] Since the genomic reads obtained by sequencing have the adapter sequence used during sequencing added thereto, the adapter sequence portion was deleted as a pretreatment for generating contigs from the genomic reads in step T120. In addition, genomic reads with 80% or more bases with a quality score of less than 20 were removed as genomic reads with low quality scores, and short genomic reads of 30 bp or less were removed. Furthermore, bases with a quality score of less than 20 were removed from the 3' end of each genomic read. De novo assembly was performed using the genomic reads after pretreatment as described above to generate contigs. Analysis of the genomic reads was performed using the next-generation sequence analysis software CLC Genomics Workbench (QIAGEN), information acquisition from external databases using UNIX commands (UNIX is a registered trademark), and Excel.

[0036] A BLAST search (hits were determined to be sequences with an E value of 0.0001 or less) was performed using contigs with a length of 10,000 bases or more. As a result, there were almost no sequences that matched bony fish, i.e., fish belonging to the class Osteichthyes, which are the same species as the Japanese bitterling. Therefore, 268 contigs with an average coverage of 5000 or more at the time of de novo assembly were extracted from the above contigs, and a BLAST search was performed. 212 of these hit sequences matched bony fish genome sequences.

[0037] Figure 3 is an explanatory diagram showing the results of a BLAST search for 212 contigs that hit the genome sequence of bony fish. Among the bony fish shown in Figure 3, the bony fish including Cyprinus carpio, which is the organism with the largest number of hit contigs, was selected and further analyzed as a candidate for closely related organisms to be used in step T130. The genome sequences of the candidates for closely related organisms described above are also referred to as "candidate sequences" below.

[0038] <Obtaining nucleic acid reads by RNA-seq analysis of the body surface of the bitterling: Step T140> RNA was collected as biological nucleic acid from the body surface of the first specified organism, the Japanese bitterling. Specifically, the surface of the Japanese bitterling was swabbed several times with a cotton swab, and the collected material was suspended in 1 mL of DNA / RNA Shield solution and frozen at -80°C. Proteinase K (Takara Bio Inc.), a proteolytic enzyme, was added to the thawed suspension in an amount of 1 / 40 of the suspension, and the suspension was kept at 55°C for 15 minutes to decompose mucin, which is abundant in mucus, in order to improve the efficiency of RNA extraction. Total RNA was then extracted using NucleoSpin RNA Plus and Quick-RNA Viral Kit (Zymo Research) according to the protocol. The RNA concentration in the extract was measured by fluorometric quantification using QuantiFluor RNA System (Promega). In addition, the molecular weight of RNA was estimated and the quality was evaluated by electrophoretic analysis using TapeStation (Agilent). The quality of RNA was evaluated based on the RIN value (RNA Integrity Number).

[0039] After decomposing and removing ribosomal RNA from the extracted RNA, RNA-seq analysis was performed. Sequencing of nucleic acid reads (RNA-seq reads) by next-generation sequencing (NGS) was outsourced (requested from Seibu Giken Co., Ltd.). Analysis of RNA-seq reads was performed using the next-generation sequence analysis software CLC Genomics Workbench (QIAGEN), information retrieval from external databases using UNIX commands, and Excel.

[0040] <Examination of closely related organisms for mapping: Step T150> Among the bony fish shown in Figure 3, Cyprinus carpio, which is the organism with the most hit contigs (closest related organism), was included, and eight of the 10 bony fish with two or more hit contigs were selected as candidates for closely related organisms to be used in step T130. Specifically, among the above 10 bony fish, those that did not contain genome sequences (nuclear genome sequences) in the sequences read from the database and only contained mitochondrial DNA sequences (Megalobrama pellegrini and Rhodeus ocellatus) were excluded from the candidates for closely related organisms. Using the individual genome sequences of each of these candidate closely related organisms as reference sequences, the RNA-seq reads of the surface of the Japanese bitterling were mapped.

[0041] Figure 4 shows the results of mapping the RNA-seq reads of the body surface of the bitterling, using the genome sequences of the eight closely related candidate organisms mentioned above as reference sequences. Figure 4 shows the mapping rate of the RNA-seq reads of the body surface of the bitterling, along with the Japanese name / common name, Genbank ID, and genome size of each fish. The "mapping rate" refers to the ratio of the number of RNA-seq reads aligned to the reference sequence to the number of all RNA-seq reads used for mapping.

[0042] Of the eight types of bony fish shown in FIG. 4, six types of bony fish had a mapping rate of 50% or more. In FIG. 4, the mapping rates of the six types of fish are hatched. Among these six types of bony fish, the mapping rate of Cyprinus carpio, which is the most closely related organism, was the highest at 74.62%. Therefore, it can be understood that filtering to remove host sequence information at a high rate of 74.62% is possible by performing mapping in step T150 using Cyprinus carpio as a closely related organism and removing nucleic acid reads aligned on the reference sequence from the RNA-seq reads.

[0043] As shown by hatching in FIG. 4, fish with a map rate of 50% or more were further examined as closely related organisms to be used in step T150 of the nucleic acid sequence information filtering method. Of the above six types of fish, Danio kyathit and Danio aesculapii belong to the same genus (Danio genus) as Danio rerio (zebrafish). Therefore, in the subsequent examination, Danio kyathit and Danio aesculapii were excluded, and only Danio rerio was used as a representative fish of the Danio genus, and a total of four types of fish were used as candidates (candidate fish) for closely related organisms.

[0044] Figure 5 is an explanatory diagram showing the results of mapping the RNA-seq reads of the surface of the Japanese bitterling, using a combination of the genome sequences of the four candidate fishes with a mapping rate of 50% or more as a reference sequence. In Figure 5, Cyprinus carpio, Danio rerio, Pimephales promela, and Megalobrama amblycephala are labeled "A", "B", "C", and "D", respectively, among the four candidate fishes. The same symbols are used in Figures 3 and 4, and Figures 6 to 11, which will be described later.

[0045] Conditions 1 to 4 shown in FIG. 5 show the breakdown of the combined genome sequences, and the candidate fish using the genome sequence as the reference sequence are marked with a circle. In "Condition 2", the top two fish species with the highest number of hit contigs shown in FIG. 3 are combined. In "Condition 3", Pimephales promela is further combined in addition to "Condition 2", and in "Condition 4", Megalobrama amblycephala is further combined in addition to "Condition 2". In "Condition 1", all of the above four types of candidate fish are combined. Thus, in all conditions, the genome sequence of Cyprinus carpio, which is the "organism with the highest number of hit contigs (closest related organism)", is included in the reference sequence. The map rate shown in FIG. 5 is the calculated map rate rounded to two decimal places.

[0046] As shown in Fig. 5, it was confirmed that by using a plurality of candidate fish in combination as related organisms, the mapping rate is increased compared to the case of using Cyprinus carpio alone. In particular, by using the genomic sequences of three or more candidate fish as reference sequences, the mapping rate is further increased, and a mapping rate exceeding 80% can be achieved. Under condition 1 in which all four types of candidate fish were combined, the mapping rate was 87.55%. In this way, the fish species with a large number of hits of the partial sequences (contigs) obtained from the shotgun metagenomic sequences of the genomic DNA of the host organism (the first specific organism), the catfish, are identified as related organisms, and mapping is performed using the genomic sequences of the identified related organisms as reference sequences. Even if the full-length genomic sequence of the host organism is unsequenced, it was confirmed that the operation of removing host sequence information from the RNA-seq reads on the surface of the host organism (the catfish) can be efficiently performed (for example, at a rate of 80% or more of the entire RNA-seq reads).

[0047] [Examination of the amount of genomic reads required to identify related organisms] In the [Example corresponding to the first embodiment] described above, in step T110, genomic reads with a total amount of nucleic acid sequence information of 70 Gb were obtained by shotgun metagenomic sequencing, and contigs were generated in step T120 using these 70 Gb of genomic reads. Therefore, further examination was carried out on the amount of genomic reads used for contig generation in step T120.

[0048] The sequence information of the genomic reads with a total amount of nucleic acid sequence information of 70 Gb described above was divided into 5 Gb, 2 Gb, 1 Gb, and 0.2 Gb. Then, de novo assembly was performed for each of the divided genomic read groups, and contig generation in step T120 was carried out. After that, 100 contigs were extracted from each of the divided genomic read groups in descending order of average coverage, and for each of the 100 extracted contigs, a homology search (blast search) in step T130 using a database was performed in the same manner as in the [Example corresponding to the first embodiment].

[0049] 6 to 9 are explanatory diagrams showing the results of the above-mentioned BLAST search for each of the divided genome read groups. Here, the results are shown of two genome reads each divided into 5 Gb, 2 Gb, 1 Gb, and 0.2 Gb (N=2). FIG. 6 shows the results of the genome read group divided into 5 Gb, FIG. 7 shows the results of the genome read group divided into 2 Gb, FIG. 8 shows the results of the genome read group divided into 1 Gb, and FIG. 9 shows the results of the genome read group divided into 0.2 Gb. In FIG. 6 to FIG. 9, the number of contigs hit for each teleost fish in which a contig was hit as a result of the BLAST search is shown, and the teleost fish are arranged in descending order of the total number of contigs hit in two BLAST searches. In FIG. 6 to FIG. 9, teleost fishes such as Rhodeus ocellatus, which are indicated with an asterisk, are excluded from the subsequent processing using candidates of closely related organisms because only the mitochondrial genome sequence is published in the database.

[0050] In Figures 6 to 9, the top two species with the greatest number of contigs hit by the BLAST search are shown with light gray hatching, and the species with the third highest number of contigs hit by the BLAST search is shown with dark gray hatching (excluding bony fishes whose only mitochondrial genome sequences are publicly available and bony fishes in the Danio genus other than Danio rerio). As shown in Figures 6 to 9, in both of the divided genome read groups, the top two species with the greatest number of contigs hit by the BLAST search were Cyprinus carpio (A) and Danio rerio (B).

[0051] As explained using Figure 5, when generating contigs using genome reads with a total amount of nucleic acid sequence information of 70 Gb (step T120) and performing a BLAST search (step T130), and using a combination of genome sequences from three or more closely related candidate fish species as a reference sequence to map the nucleic acid reads (step T150), an extremely high mapping rate of over 80% was achieved by combining Pimephales promelas (C) or Megalobrama amblycephala (D) as closely related organisms in addition to the top two species mentioned above.

[0052] Here, in the genome read group divided into 5 Gb (Figure 6) and the genome read group divided into 2 Gb (Figure 7), the candidate fish ranked third in the total number of contigs hit by the blast search was the above-mentioned Pimephales promelas (C), but in the genome read group divided into 1 Gb (Figure 8) and the genome read group divided into 0.2 Gb (Figure 9), the candidate fish ranked third in the total number of contigs hit by the blast search was different from both Pimephales promelas (C) and Megalobrama amblycephala (D). Therefore, in the genome read group divided into 1 Gb and the genome read group divided into 0.2 Gb, when identifying closely related organisms based on the number of hits by the homology search of the generated contigs, it is possible that the desired closely related organisms cannot be identified at the top, and it is highly likely that a high map rate of more than 80% cannot be achieved, even if a combination of three closely related organisms is identified and used as a reference sequence for mapping.

[0053] From the above results, it was confirmed that by setting the total amount of sequence information of the genome reads used to generate contigs by de novo assembly in step T120 to 2 Gb or more, it is possible to sufficiently ensure the degree to which host sequence information can be removed from the nucleic acid reads obtained from the organism-derived nucleic acid recovered from the host organism as a result of mapping in step T150. In particular, when a first organism belonging to bony fish is used, it was confirmed that by setting the total amount of sequence information of the genome reads used to generate contigs by de novo assembly in step T120 to 2 Gb or more and identifying three or more closely related organisms in step T130, it is easy to increase the map rate in the mapping in step T150 to, for example, 80% or more, thereby enhancing the effect of removing host sequence information.

[0054] [Example corresponding to the second embodiment] <Examination of the second specific organism based on the mapping rate> In the above-mentioned [Example corresponding to the first embodiment], the organism from which genomic DNA is extracted in step T100 to identify closely-related organisms to be used for filtering, and the host organism from which organism-derived nucleic acid is recovered in step T140 to obtain the nucleic acid lead to be analyzed, are both the same first specified organism. In contrast, the following describes the results of an investigation into a method of recovering organism-derived nucleic acid and obtaining the nucleic acid lead to be analyzed, using a second specified organism, which is different from the first specified organism from which genomic DNA is extracted to identify closely-related organisms to be used for filtering, as the host organism.

[0055] The first specified organism was a bitterling, and genomic DNA was extracted (step T100), genomic reads were obtained (step T110), contigs were generated (step T120), and closely related organisms were identified (step T130) in the same manner as in the previously described [Example corresponding to the first embodiment]. That is, the results of identifying the genome sequences of closely related organisms were the same as those in the [Example corresponding to the first embodiment] in which the same bitterling was used as the first specified organism, and are as shown in FIG. 3.

[0056] As candidates for the second specified organism to be used as a host organism to recover nucleic acid derived from organism in step T240 in Fig. 2, freshwater fishes, such as the Japanese dace, the Japanese dace, and the rainbow trout, which grow in an environment similar to that of the Japanese bitterling, were used. Then, by a method similar to step 140 in the above-mentioned "Example corresponding to the first embodiment", nucleic acid derived from organisms was recovered from the body surface of these freshwater fishes, total RNA was extracted, and RNA-seq reads were obtained (step T240). After that, a combination of multiple closely related organisms was used as a reference sequence under conditions similar to conditions 1, 3, and 4 shown in Fig. 5, and the RNA-seq reads obtained from the Japanese bitterling, which is the first specified organism, and the freshwater fishes, which are candidates for the second specified organism, were mapped, and the mapping rate was examined (step T250).

[0057] FIG. 10 is an explanatory diagram showing the results of mapping the RNA-seq reads obtained using the various freshwater fishes as host organisms and investigating the map rate. FIG. 10 also shows the results of using the first specific organism, the Japanese bitterling, as the host organism (the results shown in FIG. 5). The map rate of the RNA-seq reads obtained using the Japanese dace as the host organism was 80% or more under conditions 1, 3, and 4. The map rate of the RNA-seq reads obtained using the Japanese dace as the host organism was 80% or more under conditions 1 and 4, and even under condition 3, it showed a map rate close to 80%. In contrast, the map rate of the RNA-seq reads obtained using the rainbow trout as the host organism was low, about 30%, regardless of conditions 1, 3, or 4. In FIG. 10 and FIG. 11 described later, the Japanese bitterling is indicated with the symbol "α", the Japanese dace is indicated with the symbol "β", the Japanese dace is indicated with the symbol "γ", and the rainbow trout is indicated with the symbol "δ".

[0058] [Setting the range of closely related organisms and second specific organisms using phylogenetic trees] Figure 11 is an explanatory diagram showing a phylogenetic tree of 39 main species of freshwater fish (family Cyprinidae, Salmonidae, Family Oryzias, Family Ayuidae, Family Anguillididae, Family Gobyidae, and Family Trimeresurus) that live in ponds and rivers, including the freshwater fish shown in Figures 5 and 10. The phylogenetic tree shown in Figure 11 was created by aligning sequences using CLC Genomics Workbench (QIAGEN) using an approximately 170-base portion of the reference sequence of the MiFish region of the 12S ribosome, and using an algorithm based on the neighbor joining method, which is known as a molecular phylogenetic tree estimation method.

[0059] In FIG. 11, the four types of candidate fish (A to D) of closely related organisms shown in FIG. 5 and the four types of candidate second specific organisms (α to δ) as host organisms whose map rates were examined in FIG. 10 are indicated by the symbols already described. As shown in FIG. 11, all of the four types of teleost fish (A to D) that were candidates of closely related organisms in FIG. 5 belonged to the same family of Cyprinidae as the Japanese bitterling (α) used as the first specific organism, which is the source of organism-derived nucleic acid (host organism). Therefore, by identifying organisms that belong to the same family as the first specific organism as closely related organisms in step T130, it is considered that a high map rate can be achieved in step T150, and host sequence information can be successfully removed from the collected organism-derived nucleic acid even if the entire genome sequence of the first specific organism, which is the host organism, has not yet been deciphered.

[0060] Also, as shown in Figure 11, among the freshwater fishes listed as candidates for the second specified organism in Figure 10, the Japanese dace (β) and the Japanese dace (γ), which showed a high mapping rate of about 80% in step T250, belonged to the same family of Cyprinidae as the Japanese bitterling (α) used as the first specified organism in step T100. In contrast, among the freshwater fishes listed as candidates for the second specified organism in Figure 10, the rainbow trout (δ), which showed a low mapping rate of about 30% in step T250, belonged to the family of Salmonidae. Therefore, by using the second specified organism, which belongs to the same family as the first specified organism from which genomic DNA was extracted to identify closely related organisms, as the source of organism-derived nucleic acid (host organism) in step T240, it is considered that the host sequence information can be successfully removed from the collected organism-derived nucleic acid, even if the entire genome sequences of the first specified organism and the second specified organism have not yet been deciphered.

[0061] For example, when using the filtering method of nucleic acid sequence information shown in FIG. 1 or FIG. 2 to collect environmental nucleic acids including biological nucleic acids from host organisms living in various environments, analyze the biological nucleic acids, and conduct environmental surveys, it is expected that it will be necessary to use an organism whose entire genome sequence has not yet been deciphered and whose species or genus is unknown as the host organism, which is the source of biological nucleic acids. In such a case, it is possible to investigate which family the host organism belongs to by other methods, but it may also be confirmed by creating a phylogenetic tree as shown in FIG. 11. In particular, for fish, the mitochondrial DNA database is extensive, and it is considered relatively easy to estimate the family based on the phylogenetic tree as shown in FIG. 11. Note that the candidate fish (A to D) that showed a high map rate in FIG. 4 and were shown to be useful as reference sequences for mapping (T150, T250) using nucleic acid reads are the same family of Cyprinidae as the first specified organism, the Japanese bitterling (α), but are not located very close to the first specified organism (α) in the phylogenetic tree of FIG. 11, but are relatively distant from it. Therefore, based on the results of Figure 11, it is considered that it is more useful to select an organism that has a greater number of hits in contigs generated using the genomic DNA of the first specified organism, rather than selecting an organism that is closer to the first specified organism on the phylogenetic tree and is considered to have a higher homology in its genomic sequence to the first specified organism as the reference sequence.

[0062] The present disclosure is not limited to the above-mentioned embodiments, etc., and can be realized in various configurations without departing from the spirit of the present disclosure. For example, the technical features in the embodiments corresponding to the technical features in each form described in the Summary of the Invention column can be appropriately replaced or combined in order to solve some or all of the above-mentioned problems or to achieve some or all of the above-mentioned effects. Furthermore, if a technical feature is not described as essential in this specification, it can be appropriately deleted.

[0063] The present disclosure can also be realized in the following forms. [Application example 1] A method for filtering nucleic acid sequence information, comprising: Extracting genomic DNA from a first specific organism whose entire genome sequence has not yet been deciphered; Using the extracted genomic DNA, a genome read is obtained, which is a nucleic acid sequence obtained by fragmenting the genomic DNA; performing de novo assembly using the genomic reads to generate a contig, which is a partial sequence of the genomic DNA of the first specific organism; performing a homology search using a database containing genome sequence information and sequence information of the contig, and identifying closely related organisms, including those belonging to the same family as the first specific organism and having the largest number of contigs that satisfy a predetermined criterion for determining that there is homology, among the organisms whose genome sequence information is stored in the database; Recovering biological nucleic acids attached to a specific organism, which is the first specific organism or a second specific organism belonging to the same family as the first specific organism and whose entire genome sequence has not yet been deciphered, or biological nucleic acids present in the body of the specific organism, from the specific organism, and obtaining nucleic acid reads whose base sequences have been read by next-generation sequencing (NGS) from the recovered biological nucleic acids; Mapping is performed using the nucleic acid reads with the genome sequence of the closely related organism as a reference sequence, and the remaining nucleic acid reads other than the nucleic acid reads aligned on the reference sequence are obtained as nucleic acid sequence information to be analyzed. A method for filtering nucleic acid sequence information. [Application example 2] The method for filtering nucleic acid sequence information according to Application Example 1, The homology search is performed using contigs among the generated contigs, the contigs having an average coverage by the genome reads equal to or greater than a predetermined reference value. A method for filtering nucleic acid sequence information. [Application example 3] The method for filtering nucleic acid sequence information according to Application Example 1 or 2, Fish are used as the specific organisms. A method for filtering nucleic acid sequence information. [Application example 4] A method for filtering nucleic acid sequence information according to any one of Application Examples 1 to 3, comprising: The de novo assembly is performed using genome reads with a total amount of nucleic acid sequence information obtained by sequencing in shotgun metagenomic analysis of 2 Gb or more. A method for filtering nucleic acid sequence information. [Application example 5] A method for filtering nucleic acid sequence information according to any one of Application Examples 1 to 4, comprising: Identifying three or more organisms as the closely related organisms in order of the number of contigs that satisfy the predetermined criteria; The genome sequences of the three or more closely related organisms identified are used as the reference sequences. A method for filtering nucleic acid sequence information. [Application Example 6] A method for filtering nucleic acid sequence information according to any one of Application Examples 1 to 5, comprising: The specific organism is a mobile organism, The nucleic acid read is obtained from a nucleic acid derived from a living organism attached to the specific living organism, or from a nucleic acid derived from a living organism contained in a substance discharged from the specific living organism. A method for filtering nucleic acid sequence information.

Claims

1. A method for filtering nucleic acid sequence information, comprising: Extracting genomic DNA from a first specific organism whose entire genome sequence has not yet been deciphered; Using the extracted genomic DNA, a genomic read is obtained, which is a nucleic acid sequence obtained by fragmenting the genomic DNA; performing de novo assembly using the genomic reads to generate a contig, which is a partial sequence of the genomic DNA of the first specific organism; performing a homology search using a database containing genome sequence information and sequence information of the contig, and identifying closely related organisms, including organisms belonging to the same family as the first specific organism and having the largest number of contigs that satisfy a predetermined criterion for determining that there is homology, among the organisms whose genome sequence information is stored in the database; Recovering biological nucleic acids attached to a specific organism, which is the first specific organism or a second specific organism belonging to the same family as the first specific organism and whose entire genome sequence has not yet been deciphered, or biological nucleic acids present in the body of the specific organism, from the specific organism, and obtaining nucleic acid reads whose base sequences have been read by next-generation sequencing (NGS) from the recovered biological nucleic acids; Mapping is performed using the nucleic acid reads with the genome sequence of the closely related organism as a reference sequence, and the remaining nucleic acid reads other than the nucleic acid reads aligned on the reference sequence are obtained as nucleic acid sequence information to be analyzed. A method for filtering nucleic acid sequence information.

2. 2. A method for filtering nucleic acid sequence information according to claim 1, comprising: The homology search is performed using contigs among the generated contigs, the contigs having an average coverage by the genome reads equal to or greater than a predetermined reference value. A method for filtering nucleic acid sequence information.

3. 2. A method for filtering nucleic acid sequence information according to claim 1, comprising: Fish are used as the specific organisms. A method for filtering nucleic acid sequence information.

4. 2. A method for filtering nucleic acid sequence information according to claim 1, comprising: The de novo assembly is performed using genome reads with a total amount of nucleic acid sequence information obtained by sequencing in shotgun metagenomic analysis of 2 Gb or more. A method for filtering nucleic acid sequence information.

5. 2. A method for filtering nucleic acid sequence information according to claim 1, comprising: Identifying three or more organisms as the closely related organisms in order of the number of contigs that satisfy the predetermined criteria; The genome sequences of the three or more closely related organisms identified are used as the reference sequences. A method for filtering nucleic acid sequence information.

6. 2. A method for filtering nucleic acid sequence information according to claim 1, comprising: The specific organism is a mobile organism, The nucleic acid read is obtained from a nucleic acid derived from a living organism attached to the specific living organism, or from a nucleic acid derived from a living organism contained in a substance discharged from the specific living organism. A method for filtering nucleic acid sequence information.

Citation Information

Patent Citations

  • Methods, devices, and systems for the analysis of complete microbial strains of complex heterogeneous communities, the determination of their functional relationships and interactions, and the identification and synthesis of modifiers of bioreactivity based thereon

    JP2020504620A

  • Novel chitosanase chi1 as well as coding gene and application thereof

    JP2021185912A