Filtering method of nucleic acid sequence information
The method addresses the challenge of removing rRNA sequences from environmental RNA samples by using NGS and known rRNA sequence information for mapping, resulting in efficient rRNA removal and enhanced RNA sequence analysis in diverse environments.
Patent Information
- Application Number
- JP2023192202
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-10
- Publication Date
- 2025-05-22
AI Technical Summary
Existing methods struggle to efficiently remove rRNA sequences from nucleic acid samples when analyzing environmental RNA, especially in environments where the type and amount of organisms present are unknown, leading to difficulties in identifying suitable biological species for genome sequence removal and potential loss of RNA sequences during experimental decomposition.
A method involving next-generation sequencing (NGS) to recover and extract biological RNA, identify types of rRNA to be removed, and use known rRNA sequence information as a reference for mapping nucleic acid reads, thereby filtering out rRNA sequences and enhancing the proportion of RNA sequences suitable for analysis.
This method allows for efficient and simple removal of rRNA sequences from nucleic acid reads, even in environments with unknown organism types, thereby increasing the proportion of RNA sequences available for analysis and reducing the processing burden.
Smart Images

Figure 2025079494000001_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates to methods for filtering nucleic acid sequence information. [Background technology]
[0002] In order to investigate the diversity of organisms living in a specific environment, work is being carried out to recover environmental nucleic acids containing biological nucleic acids from the environment to be investigated and to analyze the recovered biological nucleic acids. In particular, when ribonucleic acid (RNA) is recovered as biological nucleic acid and analyzed, it is possible to obtain analysis results that reflect the state of the biological nucleic acid at the time of collection in the environment from which the biological nucleic acid was obtained, and it is also possible to analyze the state of genes actually expressed in the environment. In order to analyze biological nucleic acids, it is desirable to recover biological nucleic acids from a relatively wide range that constitutes the specific environment, and to use the recovered biological nucleic acids to perform comprehensive analysis of various biological species. In addition, in order to efficiently perform analysis when investigating biodiversity in an environment, it is desirable to remove nucleic acid sequences that are not suitable for analysis from the recovered biological nucleic acids.
[0003] As a method for removing nucleic acid sequence information that is not suitable for analysis from a collected nucleic acid sample, for example, there is a method for removing genome sequence information of a specific biological species that is thought not to contribute to a desired analysis such as biodiversity from the nucleic acid sequence information obtained from the collected nucleic acid sample. As an example of a technique for removing genome sequence information of a specific biological species that is thought not to contribute to the analysis, a technique is known in which nucleic acid derived from an organism is collected from a human to obtain nucleic acid sequence information, known human genome sequence information is removed from the obtained nucleic acid sequence information, and the remaining nucleic acid sequence information is used to analyze nucleic acid sequence information related to pathogens, enterobacteria, etc. Specifically, for example, a technique has been proposed in which a nucleic acid sample (RNA) is collected from a patient to obtain nucleic acid sequence information, known human genome sequence information is removed from the obtained nucleic acid sequence information, and the remaining nucleic acid sequence information is used to analyze the novel coronavirus (SARS-CoV-2) (see, for example, Non-Patent Document 1). Furthermore, as a related technique, a tool (framework) for removing sequence information of human-derived DNA, which is the source of nucleic acid collection, from nucleic acid sequence information has been proposed (see, for example, Non-Patent Document 2). [Prior art documents] [Non-patent literature]
[0004] [Non-Patent Document 1] Sun Zhaoyang et al., "Clinical characteristics of the host DNA-removed metagenomic next-generation sequencing technology for tectiong SARS-CoV-2, revealing host local immune signaling and assisting genomic epidemiology", Frontiers in Immunology, 15 November 2022, DOI:10.3389 / fimmu.2022.1016440 [Non-Patent Document 2] Robert Schmieder et al., "Fast Identification and Removal of Sequence Contamination from Genomic and Metagenomic Datasets", PLoS ONE, March 2011;6(3). Summary of the Invention [Problem to be solved by the invention]
[0005] However, for example, when analyzing biological nucleic acids in an environment where the type and amount of organisms present is unknown, it is difficult to identify the biological species whose genome sequence information should be removed as a nucleic acid sequence that is not suitable for analysis. Even if the biological species whose genome sequence information should be removed from the analysis target can be identified, if the entire genome sequence of the identified biological species has not yet been deciphered, the genome sequence information cannot be removed from the sequence information of the collected biological nucleic acid. Therefore, even when analyzing biological nucleic acids in an environment where the type and amount of organisms present is unknown, a technology for efficiently removing nucleic acid sequences that are not suitable for analysis from the collected biological nucleic acid has been desired.
[0006] Another method for removing nucleic acid sequences that are not suitable for analysis from the collected biological nucleic acids is to remove rRNA that is present in large amounts in biological RNA when ribonucleic acid (RNA) is the analysis target. A method for removing rRNA sequences from the analysis target is known in which rRNA in an RNA sample collected from the environment is experimentally decomposed in advance. However, when rRNA is experimentally decomposed, the yield of sampled RNA decreases during the decomposition process, which may cause inconveniences such as making it difficult to identify the biological species from which RNA that is relatively scarce in the environment originates.
[0007] Another possible method for removing rRNA sequences from the analysis target is to remove rRNA sequence information of various species whose base sequences have been decoded and stored in a database from the sequence information of the collected nucleic acid derived from the organism. However, since the rRNA information of all species has not been decoded, it may not be possible to remove the rRNA sequence of the desired organism. In addition, when removing all the rRNA sequence information in the database from the analysis target, the processing load for data processing for removal is large and it may be difficult to adopt. Thus, when analyzing RNA derived from organisms in an environment where the type and amount of organisms present is unknown, there has been insufficient consideration on an efficient and simple method for removing rRNA sequences from the analysis target. [Means for solving the problem]
[0008] The present disclosure can be realized in the following forms. (1) According to one embodiment of the present disclosure, a method for filtering nucleic acid sequence information is provided, which comprises recovering a biological RNA-containing material containing biological RNA, extracting biological RNA from the biological RNA-containing material, obtaining nucleic acid reads whose base sequences have been read by next-generation sequencing (NGS) from the extracted biological RNA, identifying a type of rRNA to be removed from the nucleic acid sequence information to be analyzed, obtaining sequence information of the identified type of rRNA, which is rRNA sequence information of an arbitrary biological species whose rRNA sequence information is known, performing mapping using the nucleic acid reads with the obtained rRNA sequence information of the arbitrary biological species as a reference sequence, and obtaining the remaining nucleic acid reads other than the nucleic acid reads aligned on the reference sequence as the nucleic acid sequence information to be analyzed. According to the method for filtering nucleic acid sequence information of this form, when analyzing nucleic acid reads obtained from biological-derived RNA extracted from the recovered biological-derived RNA-containing material, the types of rRNA to be removed from the nucleic acid sequence information to be analyzed are specified. Then, mapping is performed using, as a reference sequence, the sequence information of the types of rRNA thus specified, which is the sequence information of rRNA of any biological species whose rRNA sequence information is known, and the nucleic acid reads that have not been aligned are obtained, thereby filtering the nucleic acid reads. Therefore, even when it is unknown what organisms are present and in what quantities in the recovered biological-derived RNA, the rRNA sequence can be efficiently removed from the nucleic acid reads to be analyzed by a simple method. By thus removing the rRNA sequence from the nucleic acid reads to be analyzed, the proportion of the RNA sequence to be analyzed can be increased in the nucleic acid reads used as the analysis target. (2) In the method for filtering nucleic acid sequence information of the above form, the types of the rRNA to be removed from the nucleic acid sequence information to be analyzed may be at least one of 16S rRNA and 23S rRNA derived from prokaryotes, mitochondria, or chloroplasts. With such a configuration, at least one of the sequences of 16S rRNA and 23S rRNA can be efficiently removed from the nucleic acid reads used as the analysis target over a wide range of biological species. (3) In the method for filtering nucleic acid sequence information of the above form, the recovery of the biological-derived RNA-containing material may be performed using at least one of environmental water, air, and soil as a recovery source. With such a configuration, when analyzing biological-derived RNA collected from at least one of environmental water, air, and soil, the removal of the rRNA sequence from the biological-derived RNA to be analyzed can be efficiently performed over a wide range of biological species. (4) In the method for filtering nucleic acid sequence information of the above embodiment, sequence information of rRNA of one arbitrary biological species may be obtained for each of the identified rRNA types. With this configuration, it is possible to efficiently remove rRNA sequence information while reducing the processing load when removing rRNA sequence information from the nucleic acid sequence information to be analyzed. The present disclosure can be realized in various forms other than those described above, for example, in the form of a method for analyzing nucleic acid derived from a living organism, a method for investigating biodiversity, a method for monitoring biodiversity, and the like. [Brief description of the drawings]
[0009] [Figure 1] 1 is a flowchart showing a method for filtering nucleic acid sequence information according to an embodiment. [Diagram 2] An explanatory diagram showing the BLAST search results and mapping rate for eight contigs. [Diagram 3] FIG. 1 is an explanatory diagram showing the results of mapping to rRNA sequences registered in GenBank. [Figure 4] FIG. 1 is an explanatory diagram showing the results of mapping to rRNA sequences registered in GenBank. [Diagram 5] An explanatory diagram showing the results of a BLAST search of filtered reads. [Figure 6] An explanatory diagram showing the homology between 16S rRNA sequences by aligning the sequences. [Figure 7] An explanatory diagram showing the homology between 23S rRNA sequences by aligning the sequences. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0010] 1 is a flowchart showing a method for filtering nucleic acid sequence information according to an embodiment of the present disclosure. The method for filtering nucleic acid sequence information according to this embodiment can be used, for example, when investigating the diversity of organisms living in an environment to be investigated by recovering RNA derived from an organism from the environment to be investigated and analyzing the recovered RNA derived from the organism. The method for filtering nucleic acid sequence information according to this embodiment is a method for removing sequence information of rRNA, which is a nucleic acid sequence that is not suitable for use as an analysis target, from the sequence information of the recovered RNA derived from an organism when analyzing the recovered RNA derived from the organism.
[0011] When performing the method for filtering nucleic acid sequence information shown in FIG. 1, first, the researcher collects biological RNA-containing material containing biological RNA from the environment to be investigated, and extracts biological RNA from the collected biological RNA-containing material (step T100). Here, the environment to be investigated may be the hydrosphere such as rivers, lakes, and oceans, or may be on land. When the environment to be investigated is the hydrosphere, environmental water can be collected from rivers, lakes, and oceans to collect biological RNA-containing material. When the environment to be investigated is on land, biological RNA-containing material may be collected by collecting air, for example, or soil.
[0012] After extracting the RNA from the organism in step T100, the researcher then obtains nucleic acid reads whose base sequences are read by next-generation sequencing (NGS) from the extracted RNA from the organism (step T110). The nucleic acid reads obtained by comprehensively determining the base sequence of RNA by next-generation sequencing are also called RNA-seq reads.
[0013] Separate from the acquisition of nucleic acid reads in Project T110, the investigator identifies the types of rRNA that the investigator wants to remove from the nucleic acid sequence information to be analyzed, and acquires the sequence information of the identified types of rRNA, which is the sequence information of rRNA of any biological species whose rRNA sequence information is known (Project T120). Here, the "type of rRNA to be removed from the nucleic acid sequence information to be analyzed" can be at least any one of 23S rRNA, 16S rRNA, and 5S rRNA when the sequence information of prokaryotic rRNA is to be removed from the nucleic acid sequence information to be analyzed. Also, the "type of rRNA to be removed from the nucleic acid sequence information to be analyzed" can be at least any one of 28S rRNA (25S rRNA when the eukaryote is a plant), 18S rRNA, 5.8S rRNA, and 5S rRNA when the sequence information of eukaryotic rRNA is to be removed from the nucleic acid sequence information to be analyzed. Further, the "type of rRNA to be removed from the nucleic acid sequence information to be analyzed" can be at least any one of 12S rRNA and 16S rRNA when the sequence information of rRNA encoded by mitochondrial DNA is to be removed from the nucleic acid sequence information to be analyzed, and can be at least any one of 23S rRNA, 16S rRNA, 4.5S rRNA, and 5S rRNA when the sequence information of rRNA encoded by plant chloroplast DNA is to be removed from the nucleic acid sequence information to be analyzed.
[0014] In step T120, the sequence information of the rRNA type specified as above is further obtained, and the sequence information of the rRNA of an arbitrary biological species whose rRNA sequence information is known is obtained. The sequence information of the rRNA can be appropriately obtained from various databases storing the sequence information of the rRNA of various biological species whose rRNA sequences have been decoded. For example, when it is desired to remove the sequence information of the rRNA of a prokaryote from the nucleic acid sequence information to be analyzed, the sequence information of the rRNA of an arbitrary prokaryote, which is at least one of the 23SrRNA, 16SrRNA, and 5SrRNA specified as above, may be obtained from a database. Furthermore, when it is desired to remove the sequence information of the rRNA of a eukaryote from the nucleic acid sequence information to be analyzed, the sequence information of the rRNA of an arbitrary eukaryote, which is at least one of the 28SrRNA (25SrRNA when the eukaryote is a plant), 18SrRNA, 5.8SrRNA, and 5SrRNA specified as above, may be obtained from a database. In addition, when it is desired to remove sequence information of rRNA encoded by mitochondrial DNA from the nucleic acid sequence information to be analyzed, it is sufficient to obtain from a database at least one of sequence information of 12SrRNA and 16SrRNA, which is sequence information of rRNA encoded by mitochondrial DNA of any eukaryote, specified as above. In addition, when it is desired to remove sequence information of rRNA encoded by chloroplast DNA of a plant from the nucleic acid sequence information to be analyzed, it is sufficient to obtain from a database at least one of sequence information of 23SrRNA, 16SrRNA, 4.5SrRNA, and 5SrRNA, which is sequence information of rRNA encoded by chloroplast DNA of any plant, specified as above. However, when removing sequence information of 16SrRNA or 23SrRNA, sequence information derived from any prokaryote may be used, sequence information derived from mitochondria of any eukaryote may be used, or sequence information derived from chloroplast of any plant may be used.
[0015] In step T120, when acquiring sequence information of the type of rRNA specified as described above, it is sufficient to acquire sequence information of rRNA of at least one or more arbitrary species for each of the specified rRNA types, and sequence information of rRNA of two or more arbitrary species may be acquired. For example, when it is desired to remove sequence information of rRNA of a prokaryote from the nucleic acid sequence information to be analyzed, when acquiring sequence information of 23SrRNA, sequence information of 23SrRNA of one arbitrary species of bacteria may be acquired, or sequence information of 23SrRNA of two or more arbitrary species of bacteria may be acquired. However, from the viewpoint of reducing the processing load of removing rRNA sequence information from the nucleic acid sequence information to be analyzed, it is preferable to acquire sequence information of rRNA of five or less arbitrary species for each of the rRNA types specified as described above, more preferable to acquire sequence information of rRNA of three or less arbitrary species of organisms, and even more preferable to acquire sequence information of rRNA of one arbitrary species of organism. Even when acquiring and using sequence information of rRNA of one arbitrary species for each of the specified rRNA types, it is possible to sufficiently obtain the effect of removing rRNA sequence information from the nucleic acid sequence information to be analyzed.
[0016] In step T120, when acquiring sequence information of the type of rRNA identified as described above, if multiple types of rRNA are identified as the types of rRNA to be removed from the nucleic acid sequence information to be analyzed, sequence information of rRNA of the same biological species may be acquired as the sequence information of all types of rRNA, or sequence information of rRNA of different biological species may be acquired for each type of rRNA. For example, when it is desired to remove sequence information of rRNA of a prokaryote from the nucleic acid sequence information to be analyzed and sequence information of 23SrRNA and 16SrRNA is acquired, sequence information of 23SrRNA and 16SrRNA of any one type of bacteria may be acquired, or sequence information of 23SrRNA of any one type of bacteria and sequence information of 16SrRNA of any type of bacteria different from this bacteria may be acquired.
[0017] When the sequence information of rRNA of an arbitrary biological species is acquired in step T120 as described above, the researcher performs mapping using the nucleic acid read acquired in step T110 with the acquired sequence information of rRNA of the arbitrary biological species as a reference sequence. Then, the remaining nucleic acid reads other than the nucleic acid reads aligned on the reference sequence are acquired as the nucleic acid sequence information to be analyzed, thereby completing the filtering of the nucleic acid sequence information (step T130). The operation of aligning the nucleic acid reads on the reference sequence can be performed using a computer program by any known method. The above mapping is performed based on the sequence homology between the reference sequence and the nucleic acid read. In step T120, when multiple types of rRNA are identified, or when sequence information of rRNA of multiple types of biological species is acquired for each identified rRNA, all of the acquired sequence information of rRNA can be used as the reference sequence, and all of the nucleic acid reads aligned on any of the reference sequences can be removed from the analysis target.
[0018] According to the method for filtering nucleic acid sequence information of the present embodiment configured as described above, when analyzing nucleic acid reads obtained from biological RNA extracted from a recovered biological RNA-containing material, the type of rRNA to be removed from the nucleic acid sequence information to be analyzed is specified, and the sequence information of the specified type of rRNA, which is the sequence information of the rRNA of an arbitrary biological species, is used as a reference sequence to perform mapping, and unaligned nucleic acid reads are obtained, thereby filtering the nucleic acid reads. Therefore, even when analyzing biological RNA in an environment where the type and amount of organisms present are unknown, the rRNA sequence can be removed from the nucleic acid read to be analyzed by an efficient and simple method. In this way, by removing the rRNA sequence from the nucleic acid read to be analyzed, the proportion of RNA sequences that are desired to be the subject of analysis can be increased in the nucleic acid reads used as the subject of analysis. In other words, by efficiently removing the sequence of rRNA that may exist in a large number of copies in biological RNA, it becomes possible to analyze nucleic acid reads that match the expression level in the analysis using biological RNA.
[0019] Since rRNA sequences are relatively highly conserved between species, it is not necessary to use rRNA sequence information of a biological species that is abundant in the environment that is the subject of analysis from which the biological RNA-containing material is collected as the rRNA sequence information used as the reference sequence for mapping to remove rRNA sequences from the subject of analysis. Filtering may be performed using rRNA sequence information of a biological species that is rarely present in the environment that is the subject of investigation, and sufficient removal efficiency can be obtained even if any rRNA is appropriately selected. Therefore, the degree of freedom in selecting rRNA can be increased when obtaining rRNA sequence information of any biological species whose rRNA sequence information is known.
[0020] In addition, when multiple types of rRNA are specified as the types of rRNA to be removed from the nucleic acid sequence information to be analyzed, the removal effect according to the size of the specified rRNA can be obtained independently. In addition, regarding the removal efficiency, an additive effect can be obtained by using multiple types of rRNA.
[0021] In rRNA, for example, 5SrRNA is about 120 bases, 12SrRNA is about 880 bases, 16SrRNA is about 1500 base pairs, and 23SrRNA is about 2900 bases. The length of the nucleic acid read obtained by next-generation sequencing (NGS) is about 75 to 250 bases, and up to about 300 bases. Considering the efficiency of mapping the nucleic acid read, it is desirable that the rRNA used as the reference sequence for mapping is sufficiently longer than the nucleic acid read, so that it is desirable to use at least one of 16SrRNA and 23SrRNA derived from prokaryotes, mitochondria, or chloroplasts as the rRNA used for filtering the analysis target. It is also desirable to use 28SrRNA (25SrRNA when the eukaryote is a plant), 18SrRNA, which is the rRNA of eukaryotes, and 12SrRNA of mitochondria.
[0022] In addition, each rRNA has a relatively highly conserved region and a relatively less conserved region. For example, in 16SrRNA, hypervariable regions V1-V9 are known that are relatively less conserved and are also used for biological classification by phylogenetic analysis, and the length of these hypervariable regions is about 30-100 bp. Other regions are highly conserved between species, and are known to be used, for example, to design universal primers in amplicon analysis. rRNAs other than 16SrRNA are also known to be used in amplicon analysis. For example, in eukaryotes, 18SrRNA is used in shellfish and nematodes, 12SrRNA is used in fish and mammals, and 16SrRNA is used in birds and arthropods for amplicon analysis. The above 12SrRNA gene is a gene having a sequence homologous to the 16SrRNA gene. In these rRNAs, regions other than some of the variable regions used in amplicon analysis are highly conserved between species to the extent that they are used in the design of universal primers. Therefore, by using rRNA sequence information as a reference sequence and mapping using nucleic acid reads of approximately 75 to 250 bases in size, at least regions with relatively high conservation can be sufficiently aligned and removed, ensuring stable removal efficiency regardless of the biological species.
[0023] As another method for removing rRNA sequences from the nucleic acid sequence to be analyzed, for example, a method is known in which rRNA in an RNA sample is experimentally decomposed and removed in advance using an RNA degrading enzyme or the like. In contrast to such a method, the nucleic acid sequence information filtering method of the present embodiment removes rRNA sequences by data processing rather than experimental manipulation, so that the decrease in yield of sampled RNA can be suppressed. As a result, it becomes possible to analyze biological species from which RNA is derived that is relatively scarce in the environment. In addition, according to the nucleic acid sequence information filtering method of the present embodiment, it is possible to easily remove rRNA sequences from both eukaryotes and prokaryotes, for example, removing rRNA from both humans and microorganisms.
[0024] As an analysis method using nucleic acid reads obtained by next-generation sequencing (NGS) from RNA derived from an organism, for example, a method is known in which a contig is generated by de novo assembly using the obtained nucleic acid reads, and a homology search is performed using the sequence information of the generated contig and a database containing genome sequence information. Therefore, by using the method for filtering nucleic acid sequence information of this embodiment, a contig is generated using nucleic acid reads from which rRNA sequence information has been removed, and a homology search can be performed between the contig and a database containing genome sequence information. In contrast to the method of this embodiment, as another method for removing rRNA sequences from a nucleic acid sequence to be analyzed, a contig is generated from the nucleic acid reads by de novo assembly prior to removing the sequence information of rRNA, and mapping is performed using the contigs with the sequence information of rRNA as a reference sequence, and the remaining contigs other than the contigs aligned on the reference sequence are used as the nucleic acid sequence information to be analyzed. However, in such a method, if a sequence other than rRNA necessary for analysis is included in the contigs aligned on the reference sequence, the necessary sequence may also be lost by removing the aligned contigs. According to the method for removing rRNA sequences of this embodiment, alignment on a reference sequence is performed at the stage of a shorter nucleic acid read, so that removal of rRNA sequences using alignment can be performed easily with a lighter processing burden, and the possibility of losing sequences necessary for analysis other than rRNA can be suppressed. According to this embodiment, the processing burden for removing rRNA sequences from the nucleic acid sequence to be analyzed is reduced, thereby reducing the overall processing burden of biodiversity surveys that involve the detection of many organisms present in the environment and handle a large amount of data, and the overall processing can be made more efficient.
[0025] In the above embodiment, the environment is the target of investigation, and at least one of environmental water, air, and soil is used as the source for recovering the biological RNA-containing material from the environment. However, a different configuration may be used. Specifically, instead of using a source in which biological RNA exists in an extremely dilute state, such as the above-mentioned environmental water, air, and soil, it is also possible to use a source in which biological RNA-containing material exists in a more concentrated state. For example, a method is considered in which an environmental moving object (including living and non-living things) to which biological RNA-containing material adheres and is concentrated by moving in the environment, or other non-mobile organisms are used as the source for recovering biological RNA-containing material. Examples of organisms that are the above-mentioned environmental moving objects include organisms such as fish that move in environmental water, and organisms such as birds and insects that move on land.
[0026] The recovery of the biological RNA-containing material from an organism can be carried out, for example, from the body surface of the organism or from a substance detached from the body surface (such as fish scales or bird feathers). Alternatively, the recovery of the biological RNA-containing material from an organism can be carried out from internal substances such as the digestive tract contents of the organism or from excrement. When recovering the biological RNA-containing material from an organism (animal) as described above, the investigation target may be the bacterial flora of the body surface or digestive tract of the animal, rather than the environment. In addition, the biological RNA-containing material may be recovered from a recovery source such as seeds or spores that are released from plants or the like into the environment and move, or a part of an organism (such as fallen leaves) that is detached from plants or the like and moves through the environment. Regardless of which recovery source is used, when a recovery source containing rRNA (particularly bacterial rRNA) as a nucleic acid sequence that is not suitable for use as an analysis target is used, the nucleic acid sequence that is not suitable for use as an analysis target can be efficiently removed from the sequence information of the recovered biological RNA by applying the method for filtering nucleic acid sequence information of the present embodiment.
[0027] As described above, when organisms such as fishes moving in environmental waters or birds or insects moving on land are used as the source of the RNA-containing material, and when the RNA-containing material is collected from the body surface of the organism, the RNA extracted from the RNA-containing material may contain a relatively large amount of RNA from the organism that is the source. In such a case, when genome sequence information related to the source is known and can be used to remove the RNA from the organism that is the source, the genome sequence information is used as a reference sequence to map the RNA-seq reads obtained in step T110, thereby making it possible to remove the RNA from the organism that is the source. Such a process may be performed in addition to the process of removing rRNA sequence information by the method for filtering nucleic acid sequence information of this embodiment. EXAMPLES
[0028] <Example 1> Narrowing down sequences to be removed from the nucleic acid sequence to be analyzed The survey was conducted on environmental water (ponds) inhabited by teleost fish such as the Japanese bitterling, and sequences that should be removed from the RNA of the organisms to be analyzed were searched for. Sequences that should be removed from the analysis targets were also searched for, including the genome sequences of teleost fish that live in the environmental water.
[0029] [RNA-seq analysis of environmental water] RNA was collected as biological nucleic acid from environmental water (pond) inhabited by teleost fish such as the Japanese bitterling. Specifically, 300 mL of environmental water was collected and filtered, and the filter with the residue containing biological nucleic acid was crushed and suspended in the liquid, and total RNA was extracted from the resulting suspension. Total RNA was extracted using NucleoSpin RNA Plus and Quick-RNA Viral Kit (Zymo Research) according to the protocol. The RNA concentration in the extract was measured by fluorometric quantification using QuantiFluor RNA System (Promega). In addition, the molecular weight of RNA was estimated and the quality was evaluated by electrophoretic analysis using TapeStation (Agilent). RNA quality was evaluated based on the RIN value (RNA Integrity Number; Wong., KS; Pang., H. 2013. Smplifying HT RNA quality & quantity analysis. Genetic Engineering and Biotechnology News. 33(2) ) of 7.0 or more, which indicates the degree of degradation.
[0030] Ribosomal RNA was degraded and removed from the extracted RNA using the MGI Easy rRNA Depletion Kit (MGI Tech Co., Ltd., for human, mouse, and rat), and then RNA-seq analysis was performed. RNA-seq sequencing by next-generation sequencing (NGS) was outsourced (requested from Seibu Giken Co., Ltd.). Analysis of RNA-seq reads was performed using the next-generation sequencing analysis software CLC Genomics Workbench (QIAGEN), information retrieval from external databases using UNIX commands (UNIX is a registered trademark), and Excel.
[0031] [Shotgun metagenomic analysis of the Japanese bitterling] The tail fin of the Japanese bitterling, a type of bony fish that lives in environmental waters, was cut into approximately 5 mm squares and suspended in 1 mL of DNA / RNA Shield solution, a preservation reagent for DNA and RNA samples. Proteinase K (Takara Bio Inc.), a proteolytic enzyme, was added to the resulting suspension in an amount of 1 / 40 of the suspension, and the suspension was incubated at 55°C for 15 minutes to decompose mucin, which is abundant in the mucus, in order to improve the efficiency of DNA extraction. DNA was then extracted using the ZymoBIOMICS DNA Miniprep Kit (Zymo Research) according to the protocol. The DNA concentration in the extract was measured by fluorometric quantification using the QuantiFluor ONE dsDNA System (Promega).
[0032] A shotgun metagenomic analysis was performed using DNA extracted from the Japanese bitterling. Shotgun metagenomic sequencing by next-generation sequencing (NGS) was outsourced (requested from Seibu Giken Co., Ltd.). Through shotgun metagenomic sequencing, genome reads with a total amount of nucleic acid sequence information of 70 Gb were obtained.
[0033] The genomic reads obtained by sequencing have the adapter sequence used during sequencing added to them, so the adapter sequence was deleted as a preprocessing step when generating contigs from the genomic reads. In addition, genomic reads with 80% or more bases with a quality score of less than 20 were removed as genomic reads with low quality scores, and short genomic reads of 30 bp or less were removed. In addition, bases with a quality score of less than 20 were removed from the 3' end of each genomic read. De novo assembly was performed using the genomic reads after the preprocessing as described above to generate contigs. As a result, 268 contigs with an average coverage of 5000 or more during de novo assembly were obtained. A blast search (E value of 0.0001 or less was determined as a hit) was performed on these 268 contigs. As a result, 213 contigs were found to be in bony fishes like the Japanese bitterling, 19 in fishes other than bony fishes, 14 in bacteria, 5 in insects, 5 in shellfish, 4 in birds, 4 in mammals, 3 in plants, and 1 in amphibians. Analysis of the genome reads was performed using the next-generation sequencing analysis software CLC Genomics Workbench (QIAGEN), information retrieval from external databases using UNIX commands, and Excel.
[0034] [Refine removal sequences] The 268 contigs with an average coverage of 5000 or more during the de novo assembly were used as reference sequences to map the RNA-seq reads obtained from the environmental water described above. As a result, the mapping rate was 69.1%. The "mapping rate" refers to the ratio of the number of RNA-seq reads aligned on the reference sequence to the number of all RNA-seq reads used for mapping. In addition, when each of the 268 contigs described above was used as a reference sequence, there were eight contigs with a mapping rate of 15% or more when the RNA-seq reads were mapped. These eight contigs were used to perform a homology search (BLAST search) with the genome sequence information of various organisms stored in the database.
[0035] FIG. 2 is an explanatory diagram showing the blast search results and the mapping rate of the RNA-seq reads of environmental water for the eight contigs with a mapping rate of 15% or more (sequences 1 to 8 in FIG. 2). As shown in FIG. 2, all of the eight contigs were hit by bacterial rRNA sequences as a result of the blast search. Of the eight contigs in FIG. 2, the mapping rate to sequence 1 (contig_8177), sequence 3 (contig_1790), and sequence 5 (contig_1107), which are contigs estimated to be 16SrRNA as a result of the blast search, was 20% or more. Furthermore, the mapping rate to sequence 2 (contig_5134) and sequence 4 (contig_7573), which are contigs estimated to be 23SrRNA, was 30% or more. Thus, it was confirmed that the contigs were mapped not only to 16SrRNA but also to 23SrRNA.
[0036] FIG. 2 further shows the result of mapping (represented as sequence 9 in FIG. 2) using as a reference sequence sequence 3 (contig_1790), a contig estimated to be 16SrRNA of Aeromonas veronii, and sequence 4 (contig_7573), a contig estimated to be 23SrRNA of Aeromonas veronii. FIG. 2 also shows the result of mapping (represented as sequence 10 in FIG. 2) using as a reference sequence sequence sequence 1 (contig_8177), a contig estimated to be 16SrRNA of Vogesella sp., and sequence 2 (contig_5134), a contig estimated to be 23SrRNA of Vogesella sp. As can be seen from Figure 2, the calculated total of the mapping rate of sequence 3 (contig_1790), a contig estimated to be 16SrRNA of Aeromonas veronii, and the mapping rate of sequence 4 (contig_7573), a contig estimated to be 23SrRNA of Aeromonas veronii, was 54.31%. In contrast, the mapping rate of the two sequences, sequence 3 (contig_1790) and sequence 4 (contig_7573), was 54.32% (sequence 9), which was almost the same as the calculated total. In addition, the calculated total of the mapping rate of sequence 1 (contig_8177), a contig estimated to be 16SrRNA of Vogesella sp., and the mapping rate of sequence 2 (contig_5134), a contig estimated to be 23SrRNA of Vogesella sp., was 57.16%. In contrast, the mapping rate for sequence 1 (contig_8177) and sequence 2 (contig_5134) was 57.16% (sequence 10), which was consistent with the total calculated value above. Therefore, it is believed that the mapping rate can be additively improved by combining sequences of different types of rRNA to use as a reference sequence.
[0037] <Example 2> Verification using registered sequences of 16SrRNA and 23SrRNA [Evaluation of Genebank sequences of rRNA hits in the eight contigs with high mapping rates] The 16SrRNA and 23SrRNA sequence information (GenBank sequence) of Aeromonas veronii that hit the contigs with high mapping rates, sequence 3 (contig_1790) and sequence 4 (contig_7573) shown in Figure 2, were extracted from the database, and the sequences were used as reference sequences to map the RNA-seq reads of environmental water as described above. As shown in Figure 2, not only the contigs estimated to be Aeromonas veronii (sequences 3 and 4) but also the contigs estimated to be Vogesella sp. (sequences 1 and 2) showed similarly high mapping rates, but since the 23SrRNA sequence of Vogesella sp. was not registered in GenBank, mapping was performed only with the GenBank sequence of Aeromonas veronii.
[0038] FIG. 3 is an explanatory diagram showing the results of mapping of RNA-seq reads of environmental water using the sequence information of 16SrRNA and 23SrRNA of Aeromonas veronii registered in GenBank as a reference sequence. As shown in FIG. 3, the mapping rate for each of the 16SrRNA and 23SrRNA sequences was 21.71% and 30.37%. In addition, when both the 16SrRNA and 23SrRNA sequences were used as the reference sequence, the mapping rate was 52.08%, which was the same as the sum of the mapping rate of 16SrRNA alone and the mapping rate of 23SrRNA alone. Therefore, it was confirmed that the 16SrRNA and 23SrRNA sequences were mapped without overlap. When the sequence information of 5SrRNA of Aeromonas veronii was used as the reference sequence, the mapping rate was almost 0%. Therefore, when the rRNA sequence of Aeromonas veronii is used, it is considered that 16SrRNA and 23SrRNA are mainly mapped.
[0039] [Evaluation of rRNA sequences from other bacteria] To evaluate the effectiveness of filtering using rRNA sequence information of other bacteria, we extracted 16SrRNA and 23SrRNA sequence information of Escherichia coli, a representative bacterium, from the database and performed verification.
[0040] Figure 4 is an explanatory diagram showing the results of mapping RNA-seq reads of environmental water using the sequence information of 16SrRNA and 23SrRNA of Escherichia coli registered in GenBank as a reference sequence. As shown in Figure 4, the mapping rate for the 16SrRNA and 23SrRNA sequences, respectively, was 23.06% and 29.02%. In addition, when both the 16SrRNA and 23SrRNA sequences were used as the reference sequence, the mapping rate was 52.09%, which was almost the same as the sum of the mapping rate of 16SrRNA alone and the mapping rate of 23SrRNA alone (52.08%). Therefore, it was confirmed that the same effect of removing rRNA from RNA-seq reads could be obtained even when the rRNA sequence of another bacterium, Escherichia coli, was used.
[0041] [Bacterial community analysis using RNA-seq reads after mapping to rRNA of Aeromonas veronii and Escherichia coli] FIG. 5 is an explanatory diagram showing the results of a BLAST search of the remaining RNA-seq reads after mapping of Aeromonas veronii and Escherichia coli using the rRNA sequences (both 16SrRNA and 23SrRNA sequences) registered in GenBank as reference sequences, as explained using FIG. 3 and FIG. 4. Specifically, FIG. 5 shows the bacterial flora composition at the bacterial genus level in the BLAST search results in a stacked bar graph. FIG. 5 also shows the names of the bacteria that rank within the top 10 in terms of percentage in the bacterial flora composition.
[0042] As shown in Figure 5, no significant difference was observed in the bacterial flora composition when either the rRNA sequence derived from Aeromonas veronii or Escherichia coli was used as the reference sequence. In addition, it can be understood from the analysis results of the bacterial flora composition in Figure 5 that both Aeromonas veronii and Escherichia coli, which were used as reference sequences in mapping, are bacteria with relatively low abundance in the environmental water from which the biological RNA was collected. Therefore, it is considered that the same effect can be obtained when removing rRNA sequences using rRNA derived from either bacteria as a reference sequence.
[0043] [Bacterial 16SrRNA and 23SrRNA sequence homology] FIG. 6 is an explanatory diagram showing the homology between the sequences by aligning the 16SrRNA sequences of Aeromonas veronii, E. coli, and Vogesella sp. Also, FIG. 7 is an explanatory diagram showing the homology between the sequences by aligning the 23SrRNA sequences of several types of bacteria in addition to Aeromonas veronii and E. coli. Specifically, FIG. 7 shows the results of alignment for Aeromonas veronii and E. coli, as well as Bacillus subtilis, Enterococcus faecalis, Lactococcus lactis, Cyanobacteria (Synechocystis sp.), and Pseudomonas aeruginosa. In FIG. 6 and FIG. 7, the homology is shown by a bar graph for each base after alignment. In Figure 6, if three bacteria share the same base, the percentage is 100%, if two bacteria share the same base, the percentage is 67%, and if three bacteria do not share the same base, the percentage is 33%. In Figure 7, if seven bacteria share the same base, the percentage is 100%, if six bacteria share the same base, the percentage is 86%, if five bacteria share the same base, the percentage is 71%, if four bacteria share the same base, the percentage is 57%, if three bacteria share the same base, the percentage is 43%, if two bacteria share the same base, the percentage is 29%, and if seven bacteria do not share the same base, the percentage is 14%.
[0044] As shown in Figures 6 and 7, both 16SrRNA and 23SrRNA are highly homologous among bacteria, so it is believed that by mapping using the 16SrRNA or 23SrRNA sequence as a reference sequence, it is possible to widely remove 16SrRNA and 23SrRNA sequences from all bacteria. Figures 6 and 7 also show that 23SrRNA has a higher overall homology than 16SrRNA.
[0045] In addition, in Figure 7, the ranges of relatively low homology are indicated by curly brackets. From the length and arrangement of such relatively low homology ranges and other relatively high homology ranges, it can be understood that by performing mapping using nucleic acid reads of about 75 to 250 bases in size with the sequence information of the rRNA as a reference sequence, at least the relatively highly conserved regions can be sufficiently aligned and removed, and stable removal efficiency can be ensured regardless of the organism species.
[0046] Example 3: Effects of rRNA from organisms other than bacteria [Evaluation of rRNA sequences of aquatic plants] The effect of filtering using rRNA sequences other than bacteria was verified. As the rRNA sequence other than bacteria, 16SrRNA derived from the chloroplasts of the aquatic plant ragweed (Nymphoides peltata) was used, and the sequence information of this 16SrRNA was used as a reference sequence to map the RNA-seq reads obtained from environmental water in Example 1. As a result, the mapping rate was 17.88%. Although the mapping rate was lower than when 16SrRNA derived from bacteria was used as shown in Example 2 (see Figures 2 to 4), it was confirmed that 16SrRNA could be widely removed by mapping, even when 16SrRNA derived from aquatic plants was used.
[0047] The present disclosure is not limited to the above-mentioned embodiments, etc., and can be realized in various configurations without departing from the spirit of the present disclosure. For example, the technical features in the embodiments corresponding to the technical features in each form described in the Summary of the Invention column can be appropriately replaced or combined in order to solve some or all of the above-mentioned problems or to achieve some or all of the above-mentioned effects. Furthermore, if a technical feature is not described as essential in this specification, it can be appropriately deleted.
[0048] The present disclosure can also be realized in the following forms. [Application example 1] A method for filtering nucleic acid sequence information, comprising: Recovering a biological RNA-containing material containing biological RNA, and extracting biological RNA from the biological RNA-containing material; Obtaining nucleic acid reads whose base sequences are read by next-generation sequencing (NGS) from the extracted RNA derived from the organism; Among the rRNAs, a type of rRNA to be removed from the nucleic acid sequence information to be analyzed is identified, and sequence information of the identified type of rRNA is obtained for an arbitrary biological species for which rRNA sequence information is known; Using the acquired sequence information of the rRNA of the arbitrary biological species as a reference sequence, mapping is performed using the nucleic acid reads, and the remaining nucleic acid reads other than the nucleic acid reads aligned on the reference sequence are acquired as nucleic acid sequence information to be analyzed. A method for filtering nucleic acid sequence information. [Application example 2] The method for filtering nucleic acid sequence information according to Application Example 1, The type of rRNA to be removed from the nucleic acid sequence information to be analyzed is at least one of 16S rRNA and 23S rRNA derived from prokaryotes, mitochondria, or chloroplasts. A method for filtering nucleic acid sequence information. [Application example 3] The method for filtering nucleic acid sequence information according to Application Example 1 or 2, The RNA-containing material derived from a living organism is collected from at least one of environmental water, air, and soil. A method for filtering nucleic acid sequence information. [Application example 4] A method for filtering nucleic acid sequence information according to any one of Application Examples 1 to 3, comprising: As sequence information of the identified rRNA types, sequence information of one type of rRNA of any biological species is obtained for each of the identified rRNA types. A method for filtering nucleic acid sequence information.
Claims
1. A method for filtering nucleic acid sequence information, comprising: Recovering a biological RNA-containing material containing biological RNA, and extracting biological RNA from the biological RNA-containing material; Obtaining nucleic acid reads whose base sequences are read by next-generation sequencing (NGS) from the extracted RNA derived from the organism; Among the rRNAs, a type of rRNA to be removed from the nucleic acid sequence information to be analyzed is identified, and sequence information of the identified type of rRNA is obtained, the sequence information being of the rRNA of an arbitrary biological species for which rRNA sequence information is known; Using the acquired sequence information of the rRNA of the arbitrary biological species as a reference sequence, mapping is performed using the nucleic acid reads, and the remaining nucleic acid reads other than the nucleic acid reads aligned on the reference sequence are acquired as nucleic acid sequence information to be analyzed. A method for filtering nucleic acid sequence information.
2. 2. A method for filtering nucleic acid sequence information according to claim 1, comprising: The type of rRNA to be removed from the nucleic acid sequence information to be analyzed is at least one of 16S rRNA and 23S rRNA derived from prokaryotes, mitochondria, or chloroplasts. A method for filtering nucleic acid sequence information.
3. 2. A method for filtering nucleic acid sequence information according to claim 1, comprising: The biological RNA-containing material is collected from at least one of environmental water, air, and soil. A method for filtering nucleic acid sequence information.
4. 2. A method for filtering nucleic acid sequence information according to claim 1, comprising: As the sequence information of the identified rRNA types, sequence information of one type of rRNA of any biological species is obtained for each of the identified rRNA types. A method for filtering nucleic acid sequence information.
Citation Information
Patent Citations
Methods, devices, and systems for the analysis of complete microbial strains of complex heterogeneous communities, the determination of their functional relationships and interactions, and the identification and synthesis of modifiers of bioreactivity based thereon
JP2020504620A
Novel chitosanase chi1 as well as coding gene and application thereof
JP2021185912A