Internal standard nucleic acid for genomic or metagenomic analysis
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-01-13
- Publication Date
- 2026-03-11
AI Technical Summary
Existing genome and metagenomic analysis methods lack accurate means to evaluate assembly quality and correct for technical biases, particularly GC bias, which affects the quantification and accuracy of microbial gene abundance.
Development of artificial nucleic acids with unique sequences not found in nature, designed to include specific gene markers and controlled GC content, enabling evaluation of assembly quality and correction of GC bias using tools like CheckM.
Enables precise evaluation of assembly quality and absolute quantification of microbial genes by correcting for GC bias, ensuring accurate analysis of microbial compositions.
Smart Images

Figure 00000040_0000 
Figure 00000040_0001 
Figure 00000040_0002
Abstract
Description
[Technical field]
[0001] The present invention relates to an internal standard nucleic acid for genome or metagenomic analysis. [Background technology]
[0002] Diverse microorganisms live in all environments, including natural environments such as soil and oceans, the intestines of animals, and human living spaces such as homes. In many cases, they are established in each environment with a unique composition, and such a collection of microorganisms is called the microbiome. To analyze the microbiome, 16S rRNA gene analysis by next-generation sequencing (NGS) or whole-genome shotgun metagenomic analysis is used. 16S rRNA gene analysis comprehensively sequences the PCR products obtained by amplifying the 16S rRNA gene in the microbiome, while whole-genome shotgun metagenomic analysis comprehensively sequences the entire genomic DNA in the microbiome, which allows comprehensive analysis of the functional genes present in the microbiome and reveals the functions of the entire microbiome.
[0003] Whole genome shotgun metagenomic analysis involves the steps of extracting whole genome DNA from the microbiota, randomly fragmenting the whole genome DNA, sequencing the fragments, assembling the resulting fragment sequences (sequence reads) into a series of continuous sequences (contigs), and mapping the reads to the genome sequence estimated by the assembly, thereby quantifying the relative abundance of genes in the microbiota. However, this quantification result is only relative and cannot estimate the absolute abundance of the detected microbial groups or functional genes. In addition, the above process involves technical bias, so such biases must be accurately understood and corrected in order to obtain correct results.
[0004] For absolute quantification and quality control, a method is known in which an exogenous nucleic acid (spike-in control) having a sequence not present in the sample is used as an internal standard to correct the measured value, and a standard nucleic acid consisting of an artificial nucleic acid sequence that does not exist in nature has been developed (Patent Document 1, Non-Patent Document 1). However, bioinformatics tools for evaluating the quality of assemblies, such as CheckM (Parks et al., Genome Research, 2015, 25(7):1043-55), usually provide accurate estimates of genome completeness and contamination based on the presence or absence of specific single-copy marker genes in the assembled contigs, and therefore cannot evaluate the quality of the assembly of standard nucleic acids that do not contain gene sequences such as those described above. In addition, it is known that GC content fluctuates the sequencing coverage and reduces the accuracy of the assembly (GC bias), and a standard nucleic acid for rigorously evaluating GC bias is also desired. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] International Publication No. 2017 / 165864 [Non-patent literature]
[0006] [Non-Patent Document 1] Hardwick et al.,2018,Nature Communications,Vol.9,Article No:3096 Summary of the Invention [Problem to be solved by the invention]
[0007] The present invention has been made with an aim to provide an internal standard nucleic acid for evaluating the quality of an assembly in genome or metagenomic analysis. [Means for solving the problem]
[0008] As a result of intensive research, the present inventors have succeeded in producing an artificial nucleic acid whose assembly quality can be precisely evaluated.
[0009] That is, according to one embodiment, the present invention relates to (1) one copy each of artificial genes encoding the following non-naturally occurring sequences (a) to (p): (a) the amino acid sequence and stop codon of SEQ ID NO: 1, (b) the amino acid sequence and stop codon of SEQ ID NO: 2, (c) the amino acid sequence and stop codon of SEQ ID NO: 3, (d) the amino acid sequence and stop codon of SEQ ID NO: 4, (e) the amino acid sequence and stop codon of SEQ ID NO: 5, (f) the amino acid sequence and stop codon of SEQ ID NO: 6, (g) the amino acid sequence and stop codon of SEQ ID NO: 7, (h) the amino acid sequence and stop codon of SEQ ID NO: 8, (i) the amino acid sequence and stop codon of SEQ ID NO: 9, (j) the amino acid sequence and stop codon of SEQ ID NO: 10, and (k) the amino acid sequence and stop codon of SEQ ID NO: 11, (l) the amino acid sequence and stop codon of SEQ ID NO: 12, (m) the amino acid sequence and stop codon of SEQ ID NO: 13, (n) the amino acid sequence and stop codon of SEQ ID NO: 14, (o) the amino acid sequence and stop codon of SEQ ID NO: 15, and (p) the amino acid sequence and stop codon of SEQ ID NO: 16; (2) an artificial intergenic sequence for linking the artificial genes, each of which independently consists of a non-naturally occurring random sequence of 10 to 60 nucleotides in length; and (3) an artificial nucleic acid sequence and / or its complementary sequence or a partial fragment sequence thereof, each of which independently consists of a non-naturally occurring random sequence of 200 to 400 nucleotides in length.
[0010] The GC content of the artificial nucleic acid sequence is preferably 30 to 60%.
[0011] The artificial nucleic acid sequence is preferably selected from the group consisting of SEQ ID NOs:17-22.
[0012] The partial fragment sequence is preferably at least 300 nucleotides in length.
[0013] Furthermore, according to one embodiment, the present invention provides a nucleic acid molecule comprising an artificial nucleic acid sequence of SEQ ID NO: 23 and / or its complementary sequence, or a partial fragment sequence thereof.
[0014] The partial fragment sequence is preferably at least 300 nucleotides in length. Effect of the Invention
[0015] According to one embodiment, the nucleic acid molecule of the present invention is composed of an artificial nucleic acid sequence that does not exist in nature, but has an artificial gene sequence that can be recognized by tools such as CheckM. Therefore, the nucleic acid molecule of the present invention enables the quality evaluation of an assembly based on the presence or absence of a single-copy marker gene, which is currently commonly used.
[0016] According to one embodiment, the nucleic acid molecule of the present invention has an artificial nucleic acid sequence in which the GC content is strictly controlled, and therefore, the nucleic acid molecule of the present invention makes it possible to strictly evaluate the effect of GC bias on assembly.
[0017] Furthermore, by using the nucleic acid molecule according to the present invention, absolute quantification of genes present in a microbiota is possible. [Brief description of the drawings]
[0018] [Figure 1] FIG. 1 is a schematic diagram showing a procedure for generating an artificial nucleic acid sequence containing an artificial CDS, using seqHMM3501 as an example. [Figure 2A] FIG. 2A is a diagram showing the layout of 16 artificial CDSs in seqHMM3501, seqHMM5001, seqHMM5002, seqHMM5003, seqHMM6001, and seqHMM04. [Figure 2B]FIG. 2B shows the GC content in seqHMM3501, seqHMM5001, seqHMM5002, seqHMM5003, seqHMM6001 and seqHMM04. [Figure 2C] FIG. 2C shows the GC content in seqRANDOM01. [Figure 2D] FIG. 2D shows the pairwise sequence identity of seqHMM5002 and seqHMM5003. [Figure 3A] FIG. 3A shows the relationship between the proportion of artificial nucleic acid sequences recovered by assembly and the coverage depth when seqHMM3501, seqHMM5001, seqHMM5002, seqHMM5003, seqHMM6001, seqHMM04, and seqRANDOM01 were individually analyzed. [Figure 3B] Figure 3B shows the relationship between the number of marker genes detected from artificial nucleic acid sequences recovered by assembly and the coverage depth when seqHMM3501, seqHMM5001, seqHMM5002, seqHMM5003, seqHMM6001, seqHMM04, and seqRANDOM01 were individually analyzed. [Figure 4] FIG. 4 shows the completeness of the assembly when an equimolar mixture of seqHMM3501, seqHMM5001, seqHMM5002, seqHMM5003, seqHMM6001, seqHMM04 and seqRANDOM01 was analyzed. [Figure 5A] FIG. 5A is a plot showing relative coverage and GC content along positions in seqHMM04 and seqRANDOM01. [Figure 5B] FIG. 5B is a scatter plot showing the relationship between relative coverage and GC content along positions in seqHMM04 and seqRANDOM01. [Figure 6]FIG. 6 is a plot showing the abundance (measured and estimated) of each artificial nucleic acid in two mixtures containing different ratios of seqHMM3501, seqHMM5001, seqHMM5002, seqHMM5003, seqHMM6001, seqHMM04, and seqRANDOM01. [Figure 7] FIG. 7 is a plot showing the relative ratio (actual and estimated) of artificial nucleic acids spiked into human fecal microbiota DNA samples to human fecal microbiota DNA. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0019] The present invention will be described in detail below, but the present invention is not limited to the embodiments described in the present specification.
[0020] According to a first embodiment, the present invention provides (1) one copy each of artificial genes encoding the following sequences (a) to (p) that do not exist in nature: (a) the amino acid sequence and stop codon of SEQ ID NO: 1, (b) the amino acid sequence and stop codon of SEQ ID NO: 2, (c) the amino acid sequence and stop codon of SEQ ID NO: 3, (d) the amino acid sequence and stop codon of SEQ ID NO: 4, (e) the amino acid sequence and stop codon of SEQ ID NO: 5, (f) the amino acid sequence and stop codon of SEQ ID NO: 6, (g) the amino acid sequence and stop codon of SEQ ID NO: 7, (h) the amino acid sequence and stop codon of SEQ ID NO: 8, (i) the amino acid sequence and stop codon of SEQ ID NO: 9, (j) the amino acid sequence and stop codon of SEQ ID NO: 10, and (k) the amino acid sequence and stop codon of SEQ ID NO: 11. (l) the amino acid sequence and stop codon of SEQ ID NO: 12, (m) the amino acid sequence and stop codon of SEQ ID NO: 13, (n) the amino acid sequence and stop codon of SEQ ID NO: 14, (o) the amino acid sequence and stop codon of SEQ ID NO: 15, and (p) the amino acid sequence and stop codon of SEQ ID NO: 16; (2) an artificial intergenic sequence for linking the artificial genes, each of which independently consists of a non-naturally occurring random sequence of 10 to 60 nucleotides in length; and (3) an artificial nucleic acid sequence and / or its complementary sequence or a partial fragment sequence thereof, each of which independently consists of a non-naturally occurring random sequence of 200 to 400 nucleotides in length.
[0021] The artificial nucleic acid sequence in the nucleic acid molecule of this embodiment contains, as component (1), one copy each of artificial genes encoding the following sequences (a) to (p): In the formula, X represents any amino acid residue. (a) MXXKIKXGDXVXVIXGKXKGXXGXVXXVXXXXXXVIVEGVXXXKKXXKXXXXXXXXGXXXXXEXPIXXSNVXXXXXXXXXXXXVXXRXXXXXXKXRXXXXXGXXI (SEQ ID NO: 1) and a stop codon (b) MXXXIXXLXXXXXXXXXXXXFXXGXXVXVXXXIXEGXXXRXQXFXGXVIXXXXXGXXXXXXVXKXXXGXGVERXFXXXXXXIXXIXVXXXGXVXRAXLXYLRXXXGKXXKIKXXX (SEQ ID NO: 2) and a stop codon (c) MMAXXXRXXRVXXXIXXXIXXXLXXXIXDXXXXXXXVXXVEXSXDLXXXXVFVXXLXDXXXXXXXVXXLXXAXGFIXXXLXXXXXLXXXPXLXFXXDXSLXXXXRIXXLIXXLXXX (SEQ ID NO: 3) and a stop codon (d)MXXXFXXXPLXXGXGXTLGXXLRRVLLXXIXGXAIXXXXIXXXXXEFXXXXGVXEDVXXIIXNLKXLXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXVVXXXXXX PVXXVXy (e)MFXDXX SXLXXXXXXXX (f)MXXVAILGXXNXGKSTLLNXLXXXXXXIXSXXXXTTXXXIXGXXXXXXXQXIFIDTPGLXXXKXXXXXLLXKXIXXALXXVDLILFVVXXXXXXXXDXXLXXXLXXXXXXXXLXXXXXXXXXXXXXXXXXXXXXXXX IVXIXXXXXXXXXXXXXXXXXXLXXXXXXXPXDXVXDXXXXFXIXEXIREKILXXXXXEIPYXVXVXIXXXXXXXXXXXXIXXXIXVXRXSQKXIIIGXXGXXIKXIGXXXRXXLXXXXXXXVXLXLXVK (SEQ ID NO: 6) and stop codon (g) MXXPKXXXXXKXXXXXXXGXXXXXXXVXFGXYXLXXXXXXXIXXXXIXXXXXALXRXVXXXXXLWXRIXXXXXXXXKPXXXRMGXGKGXXEXWXXXVXXGXVLFELXGVXXXXXXXALXXAXXKLPX (SEQ ID NO: 7) and a stop codon (h) MXLLVAVSGGXDSXXLLXXLXXXXXXXXXXXXAAXVDHXXRXXSXXXXXXVXXXXXXXXXXXXXXXXXXXXXXXXXXXXXARXXRYXXLXXXXXXXXILTAHHXDDXIETILXXLXRGXXXXGLXGLXXXXXXXXXXXIXRPLLXXXKXEIXXXXXXXXLXXXXDXTNXXXXYXRNXIRXXLLP (SEQ ID NO: 8) and a stop codon (i) MINXXIXXXEVXXIXXXGXXXXIXXXXEALXXAXXXXLDLVXISXXXXXPVXKILDYGKYXYXXXKXXKXXKKXQXXIXVKEVXLXXXIXXXDXXXKXXXXXXFLXXGXXVKXXVXXXGRXXXXXXLXXXVLXXVXXXXXXXXXXXXXXXXXXXXXLLXPXXX (SEQ ID NO: 9) and a stop codon (j)MXVXLXXLXXXXXXXGXXXXXXXPXXXXFIXXXRXXXXXIXLXXXXXXLXXXXXXVXXXXXXXXXILFVGTKXXXXXXVXXXAXXXXXXYVXXRWLGGXLXNXXTIXXXIXXLXXLXXX XXXXXXXXXXXKKEXXXXXXXXXXLXXXLXGIXXLXXXPXXLXVXDXXXEXXAVXEAXXLXIPVVAXXDXNXXPXXVDXXIPXNXXXXXXXXLXXXXXXXXVXXXXXX (SEQ ID NO: 10) and stop codon (k) MXXLXLXXXDXXXXXXXNXXYRXXDXXTDVLSFXXXXXXXXXXXXXXGDLXISXXXVXXXAXXXXXXXXXXLXXHGXLHLXGYDHXXXXXXXXMXXXEXXILXXXX (SEQ ID NO: 11) and a stop codon (l) MXXXXXXXXXXXRXWXXVDAXXXXLGRLAXXVAXXLXGKXKXXYXPXXDXGDXVIVINAXXVXLXGXKXXXKXYXXXSXXXGXXXXXXXXXLXXXXXXXXLXXAVXGXLPXXXLXXXXXXXLXVYXGXXXXXXAXXPXXXXX (SEQ ID NO: 12) and a stop codon (m) MXXXKXXRXXXXRXXLLRXXXXXLLXXXXIXTTXXKXXXXXXXVEXLITXAKXXXXXXXRXVXXXLXXXXXXXXLFXXIXXXYXXRXGGYTRILKXXXRXGDXAXXAXLELVD (SEQ ID NO: 13) and a stop codon (n)MXXXXXXXXVKXLRXXTXAXXXDCKXALXXXXXDLXXAXXXLRXXGXXXXAXKKXXXXAXEGXVXXXXXXXXXXLVXIXXXTDFVAXXXXFXXLXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXLXXXXAXXXEXIXVRRIXXXXXXXXXXXX XXYXHXXXRIGVLVXXXXXXXXXXXXXLAMHVAAXXPXXLXXXXVXXXXVXXXXXIXXXXXXXXXXPXXIXXXXVXGRLXKXXXXIXLXXQXFVXXXXXXVXXXLXXXXXXVXXFXXXXVGEGIXKXXXXFXXEVXXXXXX (SEQ ID NO: 14) and stop codon (o) MMKVILXEXVXXLGXXGDXXEVKXGYAXNFLIXKXXAXXXTXXXIXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXLXIXXKXXDXGXLFGXIXXXXIXDXVXXXXXXLXKXXIXLXXXXXXXXGXXXVXLXLXXEVXAXLXVXVXXX (SEQ ID NO: 15) and a stop codon (p) MXLXXXXLXXXXXXXXXVGRGXGSGXGXTXGXGXKGXXARXXXXXXXXFEGGXXPLXXRLPXXGXXXXXXXXVXVXXXXXXXXXVXXXXLXXXXXIXXXXXXVKVLXXXXXXXXXXXXXXXXXXXXXXXXXXXXX (SEQ ID NO: 16) and a stop codon
[0022] Hereinafter, the artificial genes encoding sequences (a) to (p) are referred to as "artificial genes (a) to (p)."
[0023] The artificial genes (a) to (p) in this embodiment may be composed of any coding sequence as long as the conserved amino acids are maintained, but it is preferable to take into consideration the codon bias, homopolymer length and GC content in prokaryotes. In addition, when there are multiple codons corresponding to the conserved amino acids, any codon may be selected, and an appropriate codon may be selected taking into consideration the codon bias, homopolymer length and GC content in prokaryotes. Similarly, the stop codon may be any of the ochre codon (TAA), amber codon (TAG) and opal codon (TGA), but it is preferable to use the ochre codon in consideration of the codon bias in prokaryotes.
[0024] In the nucleic acid molecule of this embodiment, the artificial genes (a) to (p) may be arranged in any order, for example, in the 5' to 3' direction, in alphabetical order such as artificial gene (a), artificial gene (b), artificial gene (c), or in random order such as artificial gene (f), artificial gene (a), artificial gene (k).
[0025] The artificial nucleic acid sequence in the nucleic acid molecule of this embodiment includes, as a component (2), an artificial intergenic sequence for linking the artificial genes (a) to (p). The artificial intergenic sequence is a random sequence that does not exist in nature and is 10 to 60 nucleotides long, preferably 30 to 50 nucleotides long. The artificial intergenic sequence is independently composed of a random sequence for each intergenic region, and may have different lengths. Although the artificial intergenic sequence is random, it is preferable that the homopolymer length and GC content are taken into consideration.
[0026] The artificial nucleic acid sequence in the nucleic acid molecule of this embodiment includes a leading spacer sequence and a terminal spacer sequence as components (3). Specifically, a leading spacer sequence is added upstream of the artificial genes (a) to (p) linked by the artificial intergenic sequence, and a terminal spacer sequence is added downstream. The leading spacer sequence and the terminal spacer sequence are composed of random sequences having a length of 200 to 400 nucleotides, preferably 250 to 300 nucleotides, that do not exist in nature. The leading spacer sequence and the terminal spacer sequence are each independently composed of a random sequence, and may each have a different length. The leading spacer sequence and the terminal spacer sequence are random, but it is preferable that the homopolymer length and GC content are taken into consideration.
[0027] The GC content of the artificial nucleic acid sequence consisting of the above components (1) to (3) is preferably 30 to 60%. In this case, the GC content may be consistent over the entire length of the artificial nucleic acid sequence, or may vary. For example, the artificial nucleic acid sequence may have a GC content of about 30% over the entire length, or may have a region with a GC content of about 30% and a region with a GC content of about 60%.
[0028] Specific preferred examples of the artificial nucleic acid sequence include the nucleic acid sequences of SEQ ID NOs: 17 to 22. The nucleic acid sequences of SEQ ID NOs: 17 to 22 contain artificial genes (a) to (p) in alphabetical order from 5' to 3', contain 42-nucleotide artificial intergenic sequences between each gene that are independently composed of random sequences, and contain 271-nucleotide front spacer sequences and terminal spacer sequences that are independently composed of random sequences.
[0029] The nucleic acid molecule of this embodiment comprises the above artificial nucleic acid sequence and / or its complementary sequence. That is, the nucleic acid molecule of this embodiment may be either single-stranded or double-stranded. In addition, the nucleic acid molecule of this embodiment is preferably composed of DNA, but may contain a modified nucleic acid of about 1 to 3 base pairs at, for example, the end.
[0030] The nucleic acid molecule of this embodiment may include the full length of the artificial nucleic acid sequence and / or its complementary sequence, or may include a partial fragment sequence. The partial fragment sequence may be, for example, at least 300 nucleotides long, preferably 1,000 nucleotides long or more, more preferably 3,000 nucleotides long or more. In other words, it is preferable that the partial fragment sequence includes, for example, at least one artificial gene, preferably 5 or more, more preferably 8 or more artificial genes.
[0031] The nucleic acid molecule of this embodiment may be a nucleic acid consisting of only an artificial nucleic acid sequence and / or its complementary sequence or a partial fragment sequence thereof, or may be a nucleic acid consisting of only an artificial nucleic acid sequence and / or its complementary sequence or a partial fragment sequence thereof cloned into a vector. The vector that can be used in this embodiment is not particularly limited, and may be, for example, a plasmid vector such as pUC19, pT7Blue, and pGEM, a fosmid vector, a BAC vector, etc.
[0032] The nucleic acid molecule of this embodiment can be easily prepared by any conventionally known nucleic acid synthesis method.
[0033] The nucleic acid molecule of this embodiment may be added to the sample to be analyzed at an appropriate timing. For example, the nucleic acid molecule of this embodiment may be added to the sample before nucleic acid extraction, in which case, the accuracy of the entire analysis from genome DNA extraction to assembly can be controlled. Alternatively, the nucleic acid molecule of this embodiment can be added to a nucleic acid solution extracted from a microbiota sample, in which case, the quality of the assembly alone can be evaluated. A specific type of the nucleic acid molecule of this embodiment or multiple types with different sequences can be added to the sample in combination.
[0034] The sample to be analyzed may include any cell, tissue, microbiota, etc., but preferably includes a microbiota. A microbiota is a collection of multiple microorganisms present in a certain environment, and may be composed of, for example, at least 100, 300, 500, 700, 1,000, or more types of microorganisms. The types of microorganisms that constitute the microbiota are not particularly limited, and may be any classification of microorganisms, such as bacteria, fungi, protozoa, and viruses, and may include not only known microorganisms but also unknown microorganisms.
[0035] The nucleic acid molecule of this embodiment is a standard nucleic acid that is compatible with general assembly performance evaluation tools based on single-copy marker gene information, such as CheckM, and is useful for precise assembly performance evaluation in (metagenomic) analysis.
[0036] According to a second embodiment, the present invention relates to a nucleic acid molecule comprising an artificial nucleic acid sequence of SEQ ID NO: 23 and / or its complementary sequence or a partial fragment sequence thereof.
[0037] The nucleic acid molecule of this embodiment may be either single-stranded or double-stranded, similar to the nucleic acid molecule of the first embodiment. In addition, the nucleic acid molecule of this embodiment is preferably composed of DNA, similar to the nucleic acid molecule of the first embodiment, but may contain a modified nucleic acid of about 1 to 3 base pairs at, for example, the end.
[0038] The nucleic acid molecule of this embodiment may be a nucleic acid consisting of only the full length or partial fragment sequence of the artificial nucleic acid sequence and / or its complementary sequence, as in the nucleic acid molecule of the first embodiment, or may be a nucleic acid cloned into a vector. The partial fragment sequence may be, for example, at least 300 nucleotides long, preferably 1,000 nucleotides long or more, more preferably 3,000 nucleotides long or more.
[0039] The nucleic acid molecule of this embodiment may be prepared and used in the same manner as the nucleic acid molecule of the first embodiment.
[0040] The nucleic acid molecule of this embodiment is a standard nucleic acid having an artificial nucleic acid sequence with a strictly controlled GC content, and is useful for precise evaluation of assembly performance in (met)genomic analysis, in particular for evaluation of the effect of GC bias on assembly performance. EXAMPLES
[0041] The present invention will be further described below with reference to examples, but the present invention is not limited thereto.
[0042] <1. Design and synthesis of artificial nucleic acid sequences> (1-1) Design of artificial nucleic acid sequences (SEQ ID NOs: 17 to 22) containing artificial CDS Bioinformatics tools such as CheckM use a set of genes (single-copy genes) that are universal in prokaryotes and exist in only one copy per genome as markers, and evaluate the quality of the assembly based on the presence or absence of the markers in the predicted genome sequence. Therefore, in this embodiment, artificial coding sequences (CDSs) that can be recognized by general gene prediction algorithms such as Prodigal (Hyatt et al., BMC Bioinformatics, 2010, 11:119) were generated from the 16 types of marker genes shown in Table 1 below.
[0043] Table 1. Marker genes used to generate artificial CDS [Table 1]
[0044] Conserved amino acid residues in the consensus sequence extracted from each marker gene based on a hidden Markov model (HMM) were searched for and reverse translated into the corresponding DNA sequence (three-nucleotide codon). The remaining parts of each marker gene were replaced with DNA sequences encoding random amino acid residues, combined with DNA sequences encoding conserved amino acid residues, and an initiation codon (ATG) and a stop codon (TAA) were added to obtain an artificial CDS. An artificial nucleic acid sequence of 10 k nucleotides in length was generated by linking the artificial CDS with random DNA sequences (intergenic regions). An outline of the procedure for generating an artificial nucleic acid sequence is shown in Figure 1.
[0045] Six artificial nucleic acid sequences, seqHMM3501, seqHMM5001, seqHMM5002, seqHMM5003, seqHMM6001 and seqHMM04, were generated, which have the same artificial CDS order and conserved amino acid residues encoded by the artificial CDS (see SEQ ID NOs: 1 to 16), but differ in other parts (random sequences).
[0046] seqHMM3501 (SEQ ID NO: 17) [ka] [ka] [ka] [ka]
[0047] seqHMM5001 (SEQ ID NO: 18) [ka] [ka] [ka] [ka]
[0048] seqHMM5002 (SEQ ID NO: 19) [ka] [ka] [ka] [ka]
[0049] seqHMM5003 (SEQ ID NO: 20) [ka] [ka] [ka] [ka]
[0050] seqHMM6001 (SEQ ID NO: 21) [ka] [ka] [ka] [ka]
[0051] seqHMM04 (SEQ ID NO:22) [ka] [ka] [ka] [ka]
[0052] The layout of the 16 artificial CDSs in the artificial nucleic acid sequence is shown in Figure 2A, and the GC content in each sequence is shown in Figure 2B. seqHMM04 was designed to have different GC contents in each region. seqHMM5002 and seqHMM5003 were designed to contain regions with varying sequence similarity to each other to mimic sequence heterogeneity between closely related species. The pairwise sequence identity of seqHMM5002 and seqHMM5003 is shown in Figure 2D.
[0053] (1-2) Design of an artificial nucleic acid sequence (SEQ ID NO: 23) with strictly controlled GC content To accurately evaluate the effect of GC bias in the assembly, we generated an artificial nucleic acid sequence, seqRANDOM01, which is composed of completely random sequences without artificial CDS and has a strictly controlled GC content. The GC content of the artificial nucleic acid sequence, seqRANDOM01, is shown in Figure 2C.
[0054] seqRANDOM01 (SEQ ID NO:23) [ka] [ka] [ka] [ka]
[0055] It was confirmed that all of the artificial nucleic acid sequences of SEQ ID NOs: 17 to 23 have negligible similarity to base sequences registered in public databases such as NCBI (no sequences were detected that showed a similarity with an expectation value (E-value) of 0.1 or more by BLAST).
[0056] Artificial nucleic acids consisting of the sequences of SEQ ID NOs: 17 to 23 were chemically synthesized by entrusting them to GenScript Japan Co., Ltd. The artificial nucleic acids were inserted into a plasmid vector (pUC57), and the plasmid was amplified and purified by the usual procedure. The restriction enzyme sites introduced at the ends of the artificial nucleic acid sequences were cut, and the artificial nucleic acids were separated by agarose gel electrophoresis and purified.
[0057] 2. Assembly performance of artificial nucleic acids (1) Using the TruSeq DNA Nano kit (Illumina), a sequence library was individually prepared for each of the artificial nucleic acids consisting of sequences of SEQ ID NOs: 17 to 23, and sequencing was performed using the MiSeq system (Illumina) (2 × 251 bp sequencing reads). After quality control using fastp (Chen et al., Bioinformatics, 2018, 34: i884-i890), the sequencing reads were randomly sampled to vary the coverage, and assembled using the default settings of two assemblers: MEGAHIT (Li et al., Bioinformatics, 2015, 31: 1674-1676) and SPAdes (Bankevich et al., J. Comput. Biol., 2012, 19: 455-477).
[0058] The percentage of artificial nucleic acid sequences recovered by assembly is shown in Figure 3A. Both MEGAHIT (left) and SPAdes (right) results showed a sigmoidal relationship between coverage depth and assembly completeness, and also showed that complete assembly was achieved even at minimal coverage (10×).
[0059] The number of marker genes in the artificial nucleic acid sequences recovered by assembly, as detected by QUAST (Gurevich et al., Bioinformatics, 2013, 29:1072-1075) and CheckM, is shown in Figure 3B. Note that seqRANDOM01 was omitted from the CheckM analysis. All 16 genes were detected even at the minimum coverage (10x), indicating that complete assembly was achieved.
[0060] These results confirmed that the artificial nucleic acids consisting of the sequences of SEQ ID NOs: 17 to 23 are useful for evaluating the completeness of the assembly.
[0061] 3. Assembly performance of artificial nucleic acids (2) A sequencing library was prepared for an equimolar mixture of artificial nucleic acids consisting of sequences of SEQ ID NOs: 17 to 23 using a DNA Prep kit (Illumina), and sequencing was performed using the NextSeq system (Illumina) (2 × 151 bp sequencing reads). Following quality control using fastp and sampling of sequencing reads, the reads were assembled using the default settings of SPAdes.
[0062] The results are shown in Figure 4. In the figure, the shade of gray represents the sequence identity between the assembled artificial sequence and the predicted artificial sequence, and the regions with 99.9% or more identity are highlighted by a solid black line. seqHMM3501, seqHMM5001, seqHMM6001, seqHMM04 and seqRANDOM01, whose sequences are not similar to each other, were all assembled as a single contig. This result indicated that these sequences are suitable for evaluating the assembly performance. On the other hand, seqHMM5002 and seqHMM5003 coassembled due to their high sequence similarity, resulting in a fragmented assembly. This result indicated that seqHMM5002 and seqHMM5003 are useful for evaluating sequence similarity on assembly performance.
[0063] 4. Evaluation of GC bias using artificial nucleic acids For each of the artificial nucleic acids seqHMM04 (sequence number 22) and seqRANDOM01 (sequence number 23), which were designed to have different GC contents for each region, sequence libraries were prepared using the same procedure as in 3 above, and sequencing was performed. Based on the sequencing reads, coverage was calculated using BBMap (https: / / www.osti.gov / biblio / 1241166).
[0064] Plots of relative coverage (black line) and GC content (gray line) along positions in seqHMM04 and seqRANDOM01 are shown in Figure 5A. Figure 5B also presents the relative coverage and GC content in Figure 5A in a scatter plot. It was found that there was a strong correlation between sequencing coverage and GC content, and the coverage of high GC content regions was underestimated. This result indicates that seqHMM04 and seqRANDOM01 are useful for assessing GC bias.
[0065] 5. Quantitative performance of artificial nucleic acids Two types of mixtures containing artificial nucleic acids consisting of the sequences of SEQ ID NOs: 17 to 23 at different ratios were prepared, and a sequence library was prepared by the same procedure as in 3 above, sequencing was performed, and the sequencing reads were sampled.
[0066] The results are shown in Figure 6. The X-axis shows the estimated abundance (relative value) of each artificial nucleic acid, and the Y-axis shows the measured abundance (relative value) of each artificial nucleic acid. When the number of reads from the artificial nucleic acids was quantified, excellent agreement was observed between the estimated abundance and the measured abundance in all mixtures.
[0067] Next, an equimolar mixture of artificial nucleic acids consisting of sequences of SEQ ID NOs: 17 to 23 was added to the human fecal microbiota DNA sample at different mass ratios (0.3%, 1%, 3%, 31%), and a sequence library was prepared, sequencing was performed, and sequencing reads were sampled by the same procedure as in 3 above. Human fecal microbiota DNA was prepared from human feces using the ISOSPIN Fecal DNA kit (Nippon Gene Co., Ltd.) with reference to a previously published paper (Tourlousse et al., Microbiome, 2021, 9:95).
[0068] The results are shown in FIG. 7. The X-axis shows the estimated relative ratio of the artificial nucleic acid to the human fecal microbiota DNA based on the concentration calculation, and the Y-axis shows the relative ratio of the artificial nucleic acid to the human fecal microbiota DNA based on the actual measurement value. The relative ratio based on the actual measurement value was consistent with the estimated value based on the calculation. These results showed that the artificial nucleic acids consisting of the sequences of SEQ ID NOs: 17 to 23 can be used as reliable internal standards for precise absolute quantification of the amount of microorganisms.
Claims
1. (A) (1) One copy each of artificial genes encoding the following non-naturally occurring sequences (a) to (p): (a) the amino acid sequence and stop codon of SEQ ID NO: 1; (b) the amino acid sequence of SEQ ID NO:2 and a stop codon; (c) the amino acid sequence of SEQ ID NO: 3 and a stop codon; (d) the amino acid sequence and stop codon of SEQ ID NO:4; (e) the amino acid sequence and stop codon of SEQ ID NO:5; (f) the amino acid sequence and stop codon of SEQ ID NO:6; (g) the amino acid sequence and stop codon of SEQ ID NO: 7; (h) the amino acid sequence and stop codon of SEQ ID NO: 8; (i) the amino acid sequence of SEQ ID NO: 9 and a stop codon; (j) the amino acid sequence of SEQ ID NO: 10 and a stop codon; (k) the amino acid sequence of SEQ ID NO: 11 and a stop codon; (l) the amino acid sequence and stop codon of SEQ ID NO: 12; (m) the amino acid sequence and stop codon of SEQ ID NO: 13; (n) the amino acid sequence and stop codon of SEQ ID NO: 14; (o) the amino acid sequence and stop codon of SEQ ID NO: 15, and (p) the amino acid sequence of SEQ ID NO: 16 and a stop codon; (2) Artificial intergenic sequences for linking the artificial genes, each independently consisting of a random sequence of 10 to 60 nucleotides that does not occur in nature; and (3) A leading spacer sequence and a terminal spacer sequence, each independently consisting of a random sequence of 200 to 400 nucleotides that does not occur in nature. and / or a complementary sequence thereof, or (B) a partial fragment sequence of the artificial nucleic acid sequence containing at least one artificial gene encoding the sequences (a) to (p) and / or its complementary sequence A nucleic acid molecule comprising:
2. The nucleic acid molecule according to claim 1, wherein the GC content of the artificial nucleic acid sequence is 30 to 60%.
3. The nucleic acid molecule of claim 1, wherein the artificial nucleic acid sequence is selected from the group consisting of SEQ ID NOs: 17 to 22.
4. The nucleic acid molecule according to any one of claims 1 to 3, wherein the partial fragment sequence is at least 300 nucleotides in length.
5. A nucleic acid molecule comprising an artificial nucleic acid sequence of SEQ ID NO: 23 and / or its complementary sequence or a partial fragment sequence thereof.
6. The nucleic acid molecule of claim 5, wherein the subfragment sequence is at least 300 nucleotides in length.