Protein, polynucleotide, circular DNA, method for testing head and neck cancer or colorectal cancer, saliva test kit, determination marker, method for producing nucleic acid solution, and the like
Novel proteins and DNAs enhance the detection and characterization of extrachromosomal elements in human commensal bacteria, enabling accurate cancer diagnosis through saliva tests, addressing limitations in current metagenomic analyses.
Patent Information
- Application Number
- JP2025104972
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-24
- Filing Date
- 2025-06-20
- Publication Date
- 2026-01-13
AI Technical Summary
Current metagenomic analyses struggle to identify and characterize extrachromosomal elements (ECEs) in human commensal bacteria due to limited reference genome sets, missing diverse ECEs and lacking understanding of their genetic diversity and functional roles.
Development of novel proteins, polynucleotides, and circular DNAs, along with methods for testing head and neck cancer or colorectal cancer using saliva samples, including the measurement of Inocles and Streptococcus Infantis, and production of nucleic acid solutions to detect diagnostic markers.
Enhances the detection and characterization of extrachromosomal elements, providing diagnostic markers for cancer through saliva tests, improving the accuracy of cancer diagnosis and understanding of bacterial genetic diversity.
Smart Images

Figure 2026003606000037 
Figure 2026003606000038 
Figure 2026003606000039
Abstract
Description
[Technical Field]
[0001] The present invention relates to a protein, a polynucleotide, a circular DNA, a method for testing head and neck cancer or colon cancer, a saliva test kit, a diagnostic marker, a method for producing a nucleic acid solution, and the like. [Background technology]
[0002] Bacteria resident in humans are exposed to various stressors, including competition for nutrients, human drugs, antibiotics, and the human immune response. To survive these multiple stressors, bacteria expand their adaptive capabilities by acquiring accessory genes that contribute to environmental stress tolerance. Mobile genetic elements (MGEs) are the primary mechanism by which bacteria acquire accessory genes from other bacteria. ICEs (integrative conjugative elements) and transposons are intrachromosomal MGEs that transfer accessory genes to host bacteria by integrating their sequences into the chromosomes of the host bacteria (Non-Patent Documents 1 and 2). Furthermore, bacteria can acquire accessory genes via extrachromosomal elements (ECEs), which are genetic elements independent of chromosomes (Non-Patent Document 3). Microbiological studies of pathogenic bacteria have revealed that ECEs, such as plasmids, confer environmental stress tolerance, including antibiotic resistance, to host bacteria (Non-Patent Document 2). This suggests that ECEs contribute to the expansion of the adaptive capabilities of human resident bacteria.
[0003] Several advanced strategies for predicting ECEs from metagenomic sequences have revealed that human commensal bacteria possess a diverse range of ECEs, including bacteriophages (phages) (Non-Patent Documents 4-9), plasmids (Non-Patent Document 10), and viroid-like elements (Non-Patent Document 11). However, current metagenomic analyses based on limited reference genome sets may miss ECEs that cannot be classified using existing knowledge, and little is known about the genetic diversity and functional roles of ECEs in human commensal bacteria. [Prior art documents] [Non-patent literature]
[0004] [Non-Patent Document 1] Wozniak, RAF & Waldor, MK Integrative and conjugative elements: mosaic mobile genetic elements enabling dynamic lateral gene flow. Nat. Rev. Microbiol. 8, 552-563 (2010) [Non-patent document 2] Ross, K. et al. TnCentral: a Prokaryotic Transposable Element Database and Web Portal for Transposon Analysis. MBio 12, e0206021 (2021). [Non-patent document 3] Rodriguez-Beltran, J., DelaFuente, J., Leon-Sampedro, R., MacLean, RC & San Millan, A. Beyond horizontal gene transfer: the role of plasmids in bacterial evolution. Nat. Rev. Microbiol. 19, 347-359 (2021). [Non-patent document 4] Nayfach, S. et al. Metagenomic compendium of 189,680 DNA viruses from the human gut microbiome. Nature Microbiology 6, 960–970 (2021).
Direct Environment 5
Outdoor Configuration6
Direct Environment 7
Outdoor Track 8
Outdoor Tools9
[0005] An objective of the present invention is to provide novel proteins, polynucleotides, circular DNA, methods for testing head and neck cancer or colorectal cancer, saliva test kits, diagnostic markers, and methods for producing nucleic acid solutions. [Means for solving the problem]
[0006] The present inventors have conducted extensive research to solve the above problems. As a result, the present inventors have discovered a method for manufacturing a semiconductor device having the following configuration. The present invention has been completed based on the discovery that the above problems can be solved by the above method. The present invention relates to, for example, the following [1] to
[33] . [1] A protein consisting of any one of the following amino acid sequences (a) to (c): (a) an amino acid sequence represented by any one of SEQ ID NOs: 1 to 29 (b) an amino acid sequence having an identity of 90% or more but less than 100% with an amino acid sequence represented by any one of SEQ ID NOs: 1 to 29 (c) an amino acid sequence represented by any one of SEQ ID NOs: 1 to 29, in which 1 to 10 amino acids are deleted, substituted, inserted, or added; [2] A polynucleotide encoding the protein according to [1]. [3] The polynucleotide according to [2], which consists of a base sequence represented by any one of SEQ ID NOs: 30 to 58. [4] A circular DNA comprising the polynucleotide according to [2] or [3]. [5] The circular DNA according to [4], having a total length of 100 Kbp to 500 Kbp.
[0007] [6] The circular DNA according to [4] or [5], comprising a polynucleotide encoding one or more proteins selected from the group of proteins consisting of the amino acid sequences represented by SEQ ID NOs: 61 to 148 or amino acid sequences having 50% or more identity to any of the amino acid sequences represented by SEQ ID NOs: 61 to 148. [7] The circular DNA according to any one of [4] to [6], wherein the base sequence of the circular DNA is any one of the base sequences represented by SEQ ID NOs: 149 to 177. [8] A method for testing for head and neck cancer or colorectal cancer, comprising a step (C1) of measuring the amount of inocles present in the saliva of a subject. [9] The testing method according to [8], wherein the inocles are inocle-α.
[10] The testing method according to [8] or [9], wherein the step (C1) is a step of measuring the abundance of a part of the base sequence of an Inocles.
[0008]
[11] The testing method according to any one of [8] to
[10] , wherein step (C1) is a step of measuring the amount of InoC present.
[12] The testing method according to
[11] , wherein the abundance of InoC is the abundance of one or more polynucleotides selected from a group of polynucleotides consisting of the base sequences represented by SEQ ID NOs: 30 to 45.
[13] The testing method according to
[11] or
[12] , wherein the abundance of InoC is the total abundance of a group of polynucleotides consisting of the base sequences represented by SEQ ID NOs: 30 to 45.
[14] further comprising a step (C2) of measuring the amount of Streptococcus Infantis present in the saliva of the subject; and a step (C3) of calculating the amount of Inocles present per Streptcoccus Infantis by dividing the amount of Inocles present obtained in the step (C1) by the amount of Streptcoccus Infantis present obtained in the step (C2); The testing method according to any one of [8] to
[13] , comprising:
[15] The testing method according to any one of [8] to
[14] , wherein the abundance of Inocles obtained in the step (C1) or the abundance of Inocles per Streptococcus Infantis obtained in the step (C3) being smaller than a reference value indicates that the subject is highly likely to be suffering from head and neck cancer or colorectal cancer.
[0009]
[16] The testing method according to
[15] , wherein the reference value is the abundance of Inocles or the abundance of Inocles per Streptococcus Infantis obtained from a healthy subject.
[17] A saliva test kit for head and neck cancer or colorectal cancer testing, including an Inocles measurement reagent.
[18] The saliva test kit according to
[17] , further comprising a reagent for measuring Streptococcus Infantis.
[19] A marker for determining head and neck cancer or colorectal cancer, comprising the protein according to [1], the polynucleotide according to [2] or [3], or the circular DNA according to any one of [4] to [7].
[20] A marker for determining whether a patient is suffering from head and neck cancer or colorectal cancer, comprising the protein according to [1], the polynucleotide according to [2] or [3], or the circular DNA according to any one of [4] to [7].
[0010]
[21] The protein according to [1], the polynucleotide according to [2] or [3], or the circular DNA according to any one of [4] to [7], which is used as a diagnostic marker for head and neck cancer or colorectal cancer.
[22] A step (A1) of centrifuging an animal-derived specimen containing oral bacteria; and a step (A2) of treating the pellet obtained after removing the supernatant from the sample after the step (A1) with a nuclease; A method for producing a nucleic acid solution derived from oral bacteria, comprising:
[23] The manufacturing method described in
[22] , wherein the sample is saliva.
[24] The method of producing according to
[22] or
[23] , wherein the animal is a human.
[25] A primer set that specifically recognizes and amplifies at least a portion of a polynucleotide consisting of the nucleotide sequence represented by SEQ ID NOs: 149 to 177.
[0011]
[26] The primer set according to
[25] , wherein the primer set is selected from the following primer sets (A) to (E): (A) a polynucleotide consisting of the nucleotide sequence represented by SEQ ID NO: 190 and a polynucleotide consisting of the nucleotide sequence represented by SEQ ID NO: 191 (B) a polynucleotide consisting of the base sequence represented by SEQ ID NO: 192 and a polynucleotide consisting of the base sequence represented by SEQ ID NO: 193 (C) a polynucleotide consisting of the base sequence represented by SEQ ID NO: 194 and a polynucleotide consisting of the base sequence represented by SEQ ID NO: 195 (D) a polynucleotide consisting of the base sequence represented by SEQ ID NO: 196 and a polynucleotide consisting of the base sequence represented by SEQ ID NO: 197 (E) A polynucleotide consisting of the base sequence represented by SEQ ID NO: 198 and a polynucleotide consisting of the base sequence represented by SEQ ID NO: 199.
[27] The primer set according to
[25] , wherein the primer set specifically recognizes and amplifies at least a part of a polynucleotide consisting of the nucleotide sequence represented by SEQ ID NOs: 30 to 58.
[28] The primer set according to
[27] , wherein the primer set is selected from the following primer sets (Ia-A) to (Ia-D): Primer set (Ia-A): a polynucleotide consisting of the nucleotide sequence represented by SEQ ID NO: 200 and a polynucleotide consisting of the nucleotide sequence represented by SEQ ID NO: 201 Primer set (Ia-B): a polynucleotide consisting of the nucleotide sequence represented by SEQ ID NO: 202 and a polynucleotide consisting of the nucleotide sequence represented by SEQ ID NO: 203 Primer set (Ia-C): a polynucleotide consisting of the nucleotide sequence represented by SEQ ID NO: 205 and a polynucleotide consisting of the nucleotide sequence represented by SEQ ID NO: 206 Primer set (Ia-D): a polynucleotide consisting of the nucleotide sequence represented by SEQ ID NO: 208 and a polynucleotide consisting of the nucleotide sequence represented by SEQ ID NO: 209
[29] A probe capable of specifically binding to at least a part of a polynucleotide consisting of the base sequence represented by SEQ ID NOs: 149 to 177.
[30] The probe according to
[29] , which is capable of specifically binding to at least a part of a polynucleotide consisting of the base sequence represented by SEQ ID NOs: 30 to 58.
[0012]
[31] The primer set according to
[30] , wherein the probe is selected from the following probes (Ia-A) to (Ia-C): Probe (Ia-A): Polynucleotide consisting of the base sequence represented by SEQ ID NO: 204 Probe (Ia-B): Polynucleotide consisting of the base sequence represented by SEQ ID NO: 207 Probe (Ia-C): Polynucleotide consisting of the base sequence represented by SEQ ID NO: 210
[32] A method for detecting Inocles, comprising a step of hybridizing the primer set according to any one of
[25] to
[28] or the probe according to any one of
[29] to
[31] with DNA contained in saliva.
[33] A kit for detecting Inocles, comprising the primer set according to any one of
[25] to
[28] and / or the probe according to any one of
[29] to
[31] . [Effects of the Invention]
[0013] According to the present invention, it is possible to provide a protein, a polynucleotide, a circular DNA, a method for testing head and neck cancer or colon cancer, a saliva test kit, a diagnostic marker, a method for producing a nucleic acid solution, and the like. [Brief explanation of the drawings]
[0014] [Figure 1] Figure 1 shows the effect of preNuc treatment (pre-treatment with nuclease) on saliva metagenomics. [Figure 1A] Figure 1A shows a schematic workflow for high molecular weight DNA extraction from human saliva samples. Blue arrows indicate the preNuc process. [Figure 1B] Figure 1B shows the effect of preNuc treatment on human DNA removal from saliva samples. Each plot shows the percentage of human reads in NovaSeq reads obtained using preNuc-treated and untreated saliva samples. *** indicates p<0.001 based on a paired Student's t-test.
[0015] [Figure 1C]Figure 1C shows the effect of preNuc treatment on sequencing depth of microbial long reads. The bar graph shows the total number of PromethION reads and human reads per flow cell, excluding low-quality reads, obtained from preNuc-treated and untreated samples. Sample k and sample m are saliva samples from different healthy Japanese volunteers. [Figure 1D] Figure 1D shows the effect of preNuc treatment on long-read sequence length. The bar graph shows the N50 length of PromethION reads per flow cell, excluding low-quality and human reads, for preNuc-treated and untreated samples. Sample k and sample m are saliva samples from different healthy Japanese volunteers. [Figure 1E] Figure 1E is a graph of the effect of preNuc treatment on bacterial composition. The scatter plot shows the correlation between species-level bacterial composition in preNuc-treated and untreated saliva samples.
[0016] [Figure 2] Figure 2 relates to the discovery of potentially unrecognized genetic factors. [Figure 2A] Figure 2A shows the proportion of each microbial genetic element in high-quality long-read contigs. Of the 6,852 unclassified contigs, 85 were further analyzed as potential candidates for unrecognized elements. [Figure 2B] Figure 2B shows the abundance and detection rate of candidate contigs considered as unrecognized genetic factors. The scatter plot shows the average abundance (y-axis), detection rate (x-axis), and number of contigs (plot size) within each operational contig cluster obtained from 85 contigs.
[0017] [Figure 2C]Figure 2C shows the abundance of cluster 0001 contigs among all contigs in each sample. The violin plot shows the distribution of all contig abundance. The red arrows indicate the abundance of cluster 0001 contigs and their rank among all contig abundances. [Figure 2D] Figure 2D shows a box plot of contig lengths for cluster 0001. The box plot represents the interquartile range (IQR), and the line within the box indicates the median. The whiskers indicate 1.5 IQR.
[0018] [Figure 2E] Figure 2E shows a box plot of the number of ORFs (open reading frames). The box plot shows the interquartile range (IQR), and the line in the box indicates the median. The whiskers indicate 1.5 IQR. [Figure 2F] Figure 2F shows a representative replicore structure for the cluster 0001 contig (Inocle_001). The green line indicates cumulative GC skew, and the gray plot indicates GC skew at each genomic position. Candidate replication origins and termini are indicated by red (right) and blue (left) vertical lines, respectively.
[0019] [Figure 2G] Figure 2G shows a representative genome structure of Inocles, showing GC skew, GC content, and forward and reverse strands of ORFs, from the inside to the outside. [Figure 2H] Figure 2H relates to the effect of preNuc treatment on the sequencing depth of cluster 0001 contigs. The bar graph shows the average depth of cluster 0001 contigs (Inocle_004) obtained from preNuc-treated and untreated saliva samples.
[0020] [Figure 2I]Figure 2I compares the long-read contigs and corresponding short-read contigs of the cluster 0001 contig. The blue bars represent contigs constructed by long-read assembly, and the red boxes indicate insertion sequence (IS) regions. The arrows indicate fragmentation positions in the short-read assembly. The gray bar plot at the top indicates the depth of the mapped long reads. Colored bars in the depth bar plot indicate single-nucleotide variations: green for A, red for T, orange for G, and blue for C.
[0021] [Figure 3] Figure 3 relates to the identification and characterization of four Inocles taxa. [Figure 3A] Figure 3A shows the workflow for identifying additional Inocles contigs. Exploratory analysis using long-read contigs from 46 Japanese participants identified 13 Inocles contigs and the InoC gene. A representative predicted 3D structure of InoC is shown, with the green region indicating the replication-relaxation domain of InoC. To increase the number of Inocles contigs, circular contigs encoding InoC were identified using high-quality long-read contigs from 46 Japanese participants, 9 Indonesian participants, and 1 Thai participant. [Figure 3B] Figure 3B is a phylogenetic tree of the 29 Inocles amino acid sequences.
[0022] [Figure 3C] Figure 3C compares the protein repertoires of 29 Inocles. The horizontal axis represents the protein families (PFs) of Inocles, and the vertical axis represents each Inocle. The heat map shows the presence (gray) or absence (white) of each PF in each Inocle. The taxonomic groups of Inocles correspond to the color code used in Figure 3B. [Figure 3D] Figure 3D shows the detection rate of Inocles in seven countries. The size of each circle indicates the detection rate of each Inocles in each country. [Figure 3E]Figure 3E relates to the average abundance of each Inocles at each site.
[0023] [Figure 4] Figure 4 relates to the relationship between Inocles and host bacteria. [Figure 4A] Figure 4A shows the abundance of Inocles in the intracellular and extracellular particle fractions. It compares the abundance of Inocles in the bacterial pellet and the 0.45 μm filtered fraction of a saliva sample. [Figure 4B] Figure 4B shows the average number of Inocles ORFs aligned to genus-defining genes in the UniRef90 database. Error bars represent standard deviation (SD).
[0024] [Figure 4C] Figure 4C shows the experimental workflow for isolating Streptococcus strains containing Inocle_004. A saliva sample from Inocle_004 was plated on MRS agar medium. Five colonies were amplified using Inocle_004-specific primers and subjected to Sanger sequencing of the colony PCR products. Because the PCR products perfectly matched Inocle_004, three colonies were identified as Inocle_004-positive strains. Next, three Inocle_004-positive colonies (colonies 1, 2, and 3, respectively) were cultured in liquid culture in MRS medium. The absence of Inocle_004 in liquid culture colony 1 was confirmed by whole-genome sequencing. The absence of Inocle_004 in liquid culture colonies 2 and 3 was confirmed by PCR using Inocle_004-specific primers. The bacterial classification of colony 1 was determined from whole genome sequencing using GTDBtk, and the bacterial classification of colonies 2 and 3 was determined from sequencing of 16S rRNA (V1-V2 region) PCR products. The red spectrum indicates the size (bp) of the PCR product. N / A indicates unavailable.
[0025] [Figure 4D]Figure 4D shows a comparison of tetranucleotide frequencies in the chromosomes, plasmids, and phages of Inocles and Streptococcus. The heatmap shows the tetranucleotide frequencies, and a dendrogram was generated based on the similarity of tetranucleotide frequencies. [Figure 4E] Figure 4E shows a comparison of codon usage between Inoculus and Streptococcus chromosomes, plasmids, and phages. The heatmap shows the relative synonymous codon usage, and a dendrogram was generated based on the similarity of synonymous codon usage.
[0026] [Figure 5] Figure 5 relates to the functional characteristics of the Inocles. [Figure 5A] Figure 5A relates to the proportion of subcellular localization of the 1,619 Inocles protein family members. [Figure 5B] Figure 5B shows functionally annotated genes in the Inocles cytoplasmic gene cluster. In Figures 5B, 5C, and 5D, gray and white heatmaps indicate the presence of each gene in each Inocles taxon. The KEGG BRITE functional category for each gene is also shown. Genes with confirmed transcriptional activity in saliva are shown in red heatmaps.
[0027] [Figure 5C] FIG. 5C relates to genes with functional annotation of Inocles in genes with unknown subcellular localization. [Figure 5D] Figure 5D relates to functionally annotated genes of Inocles in cell wall / membrane-related genes.
[0028] [Figure 5E] FIG. 5E relates to a representative 2D structure of a bifunctional transglycosylase on a membrane. [Figure 5F] FIG. 5F relates to a representative 2D structure of sortase A on a membrane. [Figure 5G]FIG. 5G relates to a representative 2D structure of the LPXTG-like motif protein family on a membrane. [Figure 5H] FIG. 5H relates to a representative 2D structure of signal peptidase on the membrane. [Figure 5I] Figure 5I relates to the functions of homologous genes between Inocles and Streptococcus chromosomes. The bar graph shows the proportion of each gene function among all homologous genes between Inocles and Streptococcus chromosomes.
[0029] [Figure 6] FIG. 6 relates to the relationship between Inocle-α and human physiological functions. [Figure 6A] Figure 6A shows the correlation between Inocle-α and cell types in PBMCs. The heatmap shows the partial Spearman's ρ between PBMC populations and the abundance of Inocle-α and Streptococcus, with age and sex as confounding factors. * indicates p<0.05, *** indicates p<0.001.
[0030] [Figure 6B] Figure 6B shows the plasma proteome pathways that are positively correlated with Inocle-α. The bar graph shows the results of GO enrichment analysis, where the expression of proteins in plasma was significantly correlated with the abundance of Inocle-α, using partial Spearman's rho, taking into account age and sex as confounding factors.
[0031] [Figure 6C]Figure 6C compares the detection rate and abundance of Inocle-α in healthy controls (HC) and head and neck cancer patients (HNC). The bar graph shows the detection rate of Inocle-α in HC and HNC groups, and the box plot shows the abundance of Inocle-α in HC and HNC groups. The outlines in the box plots are hidden for visualization purposes. * indicates p<0.05, and *** indicates p<0.001. Statistical results were obtained using Fisher's exact test for detection rate and MaAsLin2 (considering age, sex, smoking status, and cancer treatment as confounding factors) for abundance.
[0032] [Figure 6D] Figure 6D compares the detection rate and abundance of Streptococcus in healthy controls (HC) and head and neck cancer patients (HNC). The bar graph shows the detection rate of Streptococcus in the HC and HNC groups, and the box plot shows the abundance of Streptococcus in the HC and HNC groups. The box plots hide outlines for visualization purposes. NS indicates no significant difference. Statistical results were obtained using Fisher's exact test for detection rate and MaAsLin2 (considering age, sex, smoking status, and cancer treatment as confounding factors) for abundance.
[0033] [Figure 6E] Figure 6E compares the abundance of Inocle-α and Streptococcus in patients with different diseases and HC in each study. Coefficients and p-values calculated using MaAsLin2 (taking into account age and sex as confounding factors) are shown. ** indicates p<0.01, *** indicates p<0.001. Abbreviations: CRC: colorectal cancer, PDAC: pancreatic ductal adenocarcinoma, RA: rheumatoid arthritis.
[0034] [Figure 7]Figure 7 shows the sequence similarity between long-read and short-read contigs. The sequence similarity between PromethION contigs and NovaSeq contigs was obtained from the same sample (sample IDs: m_5 and k_5) assembled using metaFlye and MEGAHIT. The y-axis shows the average sequence similarity between the best-hit NovaSeq contig and the PromethION contig corresponding to each long-read depth (x-axis) of the PromethION contig.
[0035] [Figure 8] Figure 8 relates to the workflow for classifying long-read contigs as microbial genetic elements. [Figure 8A] Figure 8A is a flowchart showing the pipeline for classifying long-read contigs as individual genetic elements. [Figure 8B] Figure 8B shows the percentage of open reading frames (ORFs) aligned to the UniRef90 database among all ORFs of each known genetic element.
[0036] [Figure 9] Figure 9 shows the long-read mapping results for 13 contigs belonging to cluster 0001. The gray bar plot displayed in the mapping results of the Integrative Genomics Viewer (IGV) indicates the depth of the mapped long reads. Forward and reverse reads are shown as pastel-colored red and blue boxes, respectively.
[0037] [Figure 10] Figure 10 shows the Replicore structure for the 13 contigs belonging to cluster 0001. The green and gray lines indicate the cumulative GC skew and the GC skew at each genomic position, respectively. The potential replication origins and endpoints are indicated by red and blue vertical lines, respectively. The blue arrows in the upper figure indicate the forward and reverse strands of the ORF.
[0038] [Figure 11]Figure 11 shows the impact of long-read metagenomics on the reconstruction of Inocles contigs. It compares the long-read contigs with the corresponding short-read contigs. The blue bars represent the complete contigs constructed by long-read assembly. The red boxes indicate insertion sequence (IS) regions within the contigs. The arrows indicate fragmented positions in the short-read assembly. The gray bar plot at the top indicates the depth of the mapped long reads. The colored bars in the depth bar plot indicate single-nucleotide variations: green for A, red for T, orange for G, and blue for C.
[0039] [Figure 12] Figure 12 shows the Replicore structure in 16 additional Inocles. The green line indicates the cumulative GC skew, and the gray plot indicates the GC skew at each genomic position. The potential replication origins and endpoints are indicated by red and blue vertical lines, respectively. The blue arrows in the upper diagram indicate the forward and reverse strands of the ORF.
[0040] [Figure 13] Figure 13 relates to genomic differences between four Inocles taxa. [Figure 13A] Figure 13A shows the amino acid identity of the InoC gene among Inocles taxa. The heat map shows the average amino acid identity (AAI) of the InoC gene in each comparison pair, and the numbers in the heat map represent the AAI. [Figure 13B] Figure 13B shows the distribution of contig sizes in each Inocles taxon. Error bars represent standard deviation (SD). In Figures 13B, 13D, 13E, and 13F, box plots represent the interquartile range (IQR), the line in the box represents the median, and the whiskers indicate 1.5 IQR.
[0041] [Figure 13C] Figure 13C shows the average nucleotide identity in each Inocles taxon. Error bars represent standard deviation (SD). [Figure 13D]Figure 13D shows the number of tRNAs in each Inocles taxon. [Figure 13E] FIG. 13E shows the GC content in each Inocles taxon. [Figure 13F] Figure 13F shows the number of insertion sequences (IS) in each Inocles taxon.
[0042] [Figure 14] FIG. 14 relates to functional annotation of cell wall-associated proteins of Inocles. [Figure 14A] Figure 14A shows the presence or absence of protein families (PFs) with LPXTG-like motifs in each Inocles taxon. Genes with confirmed transcriptional activity in saliva are shown in red heatmap. [Figure 14B] Figure 14B relates to the 2D structure of a protein family (PF) with an LPXTG-like motif on the membrane.
[0043] [Figure 14C] Figure 14C shows the presence or absence of PF containing the signal peptidase S26 domain (PF10502) in each Inocles taxon. Genes with confirmed transcriptional activity in saliva are shown in red heatmap. [Figure 14D] FIG. 14D relates to the 2D structure of signal peptidase (Inocle_PF_0398) and signal peptidase (Inocle_PF_0542) on the membrane.
[0044] [Figure 15] Figure 15 shows the number of homologous genes between the chromosomes of Inocles and Streptococcus. The heat map shows the number of homologous genes (e-value < 1e-10) between each combination of the genomes of Inocles and Streptococcus.
[0045] [Figure 16] FIG. 16 relates to the association of Inocles with gender and age. [Figure 16A]Figure 16A compares the abundance of Inocle-α in males and females. Box plots represent the interquartile range (IQR), the line in the box indicates the median, and the whiskers indicate 1.5IQR. NS indicates no significant difference based on the Wilcoxon rank sum test. [Figure 16B] FIG. 16B shows Spearman's rho between age and abundance of Inocle-α.
[0046] [Figure 17] Figure 17 compares the abundance of Inocle-α, the ratio of Inocle-α to S. infantis, and the abundance of S. infantis in healthy controls (HC) and head and neck cancer patients (HNC). Blue and red bars indicate the detection rate of Inocle-α in the HC and HNC groups, respectively. *: p<0.05, NS: not significant. Statistical results were obtained using Maaslin2 with random effects for age, sex, and smoking status. [Figure 18] Figure 18 compares the ratio of Inocle-α to S. infantis and the abundance of S. infantis between disease patients and healthy control groups. Coefficient values calculated using the Maaslin2 method are shown. *: p<0.05. Statistical results were obtained using the Maaslin2 method, using random effects for age, sex, and smoking status in the HNC study, and random effects for age and sex in the CRC study. Abbreviations: HC: healthy control group, HNC: head and neck cancer, CRC: colorectal cancer
[0047] [Figure 19] Figure 19 relates to the functional characteristics of Inocles. Schematic representation of the functional annotation of Inocles based on significant associations between gene annotations and human physiology. [Figure 20] FIG. 20 shows an electrophoretic image obtained when nucleic acid amplification of Inocle-α was performed using primer set B and the amplified product was detected by electrophoresis. DETAILED DESCRIPTION OF THE INVENTION
[0048] Next, the present invention will be described in detail. In this specification, the expression "A to B" regarding a numerical range means A or more and B or less, unless otherwise specified. Furthermore, when multiple lower limit values and multiple upper limit values are listed in the description of a certain element, the numerical range formed by combining a value arbitrarily selected from the listed lower limit value and a value arbitrarily selected from the listed upper limit value is also included. Furthermore, % means % by mass.
[0049] In this specification, when the units of the values written before and after "to" indicating a numerical range are the same, the unit of the values written before "to" may be omitted. For example, "50 mol% to 85 mol%" may be written as "50 to 85 mol%." In this specification, the amount of each component in a composition means the total amount of the multiple substances present in the composition, unless otherwise specified, when multiple substances corresponding to that component are present in the composition.
[0050] In this specification, Inocles are circular DNAs consisting of the base sequence represented by any of SEQ ID NOs: 149 to 177, and are classified into Inocle-α (SEQ ID NOs: 149 to 164, respectively) consisting of 16 types of Inocle 001 to 016, Inocle-β (SEQ ID NOs: 165 to 170, respectively) consisting of 6 types of Inocle 017 to 022, Inocle-γ (SEQ ID NOs: 171 to 174, respectively) consisting of 4 types of Inocle 023 to 026, and Inocle-δ (SEQ ID NOs: 175 to 177, respectively) consisting of 3 types of Inocle 027 to 029.
[0051] <Protein (I)> The present specification discloses a novel protein (I) consisting of any one of the following amino acid sequences (a) to (c): (a) an amino acid sequence represented by any one of SEQ ID NOs: 1 to 29 (b) an amino acid sequence having an identity of 90% or more but less than 100% with an amino acid sequence represented by any one of SEQ ID NOs: 1 to 29 (c) an amino acid sequence represented by any one of SEQ ID NOs: 1 to 29, in which 1 to 10 amino acids are deleted, substituted, inserted, or added;
[0052] (a) The amino acid sequence represented by any one of SEQ ID NOs: 1 to 29 is an amino acid sequence deduced from the gene sequence of "InoC" (Inocle conserved gene), which is a marker for Inocles. The amino acid sequences of SEQ ID NOs: 1 to 16 (amino acid sequences of Inocle-α) are amino acid sequences deduced from the gene sequence of Inocle-α (nucleotide sequences of SEQ ID NOs: 30 to 45), which was present in lower amounts in patients with head and neck cancer or colorectal cancer than in healthy individuals, as described below in the Examples. Therefore, the abundance of a protein consisting of at least one amino acid sequence selected from the group consisting of amino acid sequences represented by any of SEQ ID NOs: 1 to 16 (amino acid sequences of Inocle-α) contained in a subject's saliva can be said to be a marker indicating that the subject is likely to be suffering from head and neck cancer or colorectal cancer.
[0053] Furthermore, InoC is a sequence that characterizes Inocles, and the amino acid sequence identity between Inocles is 64.1% or more. As shown in Figure 13A, InoC has high sequence identity with Inocle-α, Inocle-β, Inocle-γ, and Inocle-δ, which are types of Inocles. Therefore, it is understood that not only InoC of Inocle-α, but also Inocle-β, Inocle-γ, and Inocle-δ InoC can be used as a diagnostic marker for head and neck cancer or colorectal cancer. The amino acid sequences of SEQ ID NOs: 17 to 22 (amino acid sequences of Inocle-β InoC), SEQ ID NOs: 23 to 26 (amino acid sequences of Inocle-γ InoC), and SEQ ID NOs: 27 to 29 (amino acid sequences of Inocle-δ InoC) share sequence identity with the amino acid sequences represented by any of SEQ ID NOs: 1 to 16 (amino acid sequences of Inocle-α InoC) as shown in Figure 13A. Because of this high identity, the presence of a protein consisting of at least one amino acid sequence selected from the group consisting of the amino acid sequences represented by any of SEQ ID NOs: 1 to 29 (amino acid sequences of Inocle-α, β, γ, δ) in the saliva of a subject can be said to be a marker indicating a high probability that the subject is suffering from head and neck cancer or colorectal cancer.
[0054] The amino acid sequences of SEQ ID NOs: 17 to 22 (amino acid sequences of Inocle-β InoC), SEQ ID NOs: 23 to 26 (amino acid sequences of Inocle-γ InoC), and SEQ ID NOs: 27 to 29 (amino acid sequences of Inocle-δ InoC) are amino acid sequences deduced from the gene sequence of Inocle-β InoC (nucleotide sequences of SEQ ID NOs: 46 to 51), the gene sequence of Inocle-γ InoC (nucleotide sequences of SEQ ID NOs: 52 to 55), and the gene sequence of Inocle-δ InoC (nucleotide sequences of SEQ ID NOs: 56 to 58), respectively.
[0055] (b) An amino acid sequence that has an identity of 90% or more but less than 100% to an amino acid sequence represented by any one of SEQ ID NOs: 1 to 29 may have an identity of, for example, 91% or more but less than 100%, 92% or more but less than 100%, 93% or more but less than 100%, 94% or more but less than 100%, 95% or more but less than 100%, 96% or more but less than 100%, 97% or more but less than 100%, 98% or more but less than 100%, or 99% or more but less than 100%. (b) An amino acid sequence that has an identity of 90% or more but less than 100% to an amino acid sequence represented by any one of SEQ ID NOs: 1 to 29 preferably has an identity of 95% or more but less than 100%, more preferably 96% or more but less than 100%, even more preferably 97% or more but less than 100%, still more preferably 98% or more but less than 100%, and particularly preferably 99% or more but less than 100%. (b) A protein consisting of an amino acid sequence that has an identity of 90% or more but less than 100% to the amino acid sequence represented by any one of SEQ ID NOs: 1 to 29 is highly identical to the protein consisting of the amino acid sequence represented by any one of SEQ ID NOs: 1 to 29, and is therefore considered to have the same function as (a) a protein consisting of the amino acid sequence represented by any one of SEQ ID NOs: 1 to 29.
[0056] The identity of amino acid sequences means the percentage (%) of matching amino acid residues relative to the length of the alignment region in optimal alignment when two amino acid sequences are aligned using a mathematical algorithm known in the art, and can be calculated using, for example, NCBI BLAST (National Center for Biotechnology Information Basic Local Alignment Search Tool).
[0057] (c) In the amino acid sequence represented by any one of SEQ ID NOs: 1 to 29, the amino acid sequence in which 1 to 10 amino acids have been deleted, substituted, inserted or added may be, for example, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids have been deleted, substituted, inserted or added. (c) In the amino acid sequence represented by any one of SEQ ID NOs: 1 to 29, in which 1 to 10 amino acids have been deleted, substituted, inserted or added, the number of deleted, substituted, inserted or added amino acids is preferably 5 or less, more preferably 4 or less, even more preferably 3 or less, even more preferably 2 or less, and particularly preferably 1. (c) A protein consisting of an amino acid sequence represented by any one of SEQ ID NOs: 1 to 29, in which 1 to 10 amino acids have been deleted, substituted, inserted or added, is highly identical to a protein consisting of an amino acid sequence represented by any one of SEQ ID NOs: 1 to 29, and is therefore considered to have the same function as (a) a protein consisting of an amino acid sequence represented by any one of SEQ ID NOs: 1 to 29.
[0058] <Polynucleotide encoding protein (I)> The present specification discloses a polynucleotide (hereinafter also referred to as polynucleotide (I)) encoding a protein (I) consisting of any one of the amino acid sequences (a) to (c) above. The base sequence of the polynucleotide (I) is not particularly limited, as long as it encodes the protein (I) consisting of any one of the amino acid sequences (a) to (c) above.
[0059] As used herein, the term "polynucleotide" refers to a chain of nucleotides of any length, including DNA and RNA. Nucleotides can be deoxyribonucleotides, ribonucleotides, modified nucleotides or bases and / or their analogs, or any substance that can be incorporated into a chain by DNA or RNA polymerase. Polynucleotides can include modified nucleotides, such as methylated nucleotides and their analogs. If present, modifications to the nucleotide structure can be imparted before or after assembly of the chain. The sequence of nucleotides can be interrupted by non-nucleotide components. Polynucleotides can be further modified after polymerization, such as by conjugation with a labeling component. Other types of modifications include, for example, "caps," substitution of one or more naturally occurring nucleotides with analogs, internucleotide modifications, such as uncharged linkages (e.g., methylphosphonate, phosphotriester, phosphoamidate, carbamate, etc.) and charged linkages (e.g., phosphorothioates, phosphorodithioates, etc.), those containing pendant moieties such as proteins (e.g., nucleases, toxins, antibodies, signal peptides, poly-L-lysine, etc.), those containing interfering substances (e.g., acridine, psoralen, etc.), those containing chelators (e.g., metals, radioactive metals, boron, metal oxides, etc.), those containing alkylating agents, those containing modified linkages (e.g., alpha-anomeric nucleic acids, etc.), as well as unmodified forms of polynucleotides. Additionally, any of the hydroxyl groups normally present on the sugar may be replaced, for example, by phosphonate groups, phosphate groups, protected by standard protecting groups, or activated to provide for additional linkages to additional nucleotides, or conjugated to a solid support. The 5' and 3' terminal OH may be phosphorylated or substituted with amines or organic capping group moieties of 1 to 20 carbon atoms. Other hydroxyls may also be derivatized to standard protecting groups.Polynucleotides may also contain analogous forms of ribose or deoxyribose sugars commonly known in the art, including, for example, 2'-O-methyl-, 2'-O-allyl, 2'-fluoro-, or 2'-azido-ribose, carbocyclic sugar analogs, alpha- or beta-anomeric sugars, epimeric sugars such as arabinose, xylose, or lyxose, pyranose sugars, furanose sugars, sedoheptulose, acyclic analogs, and abasic nucleoside analogs such as methyl riboside. One or more phosphodiester linkages may be replaced by alternative linking groups.
[0060] Polynucleotide (I) is preferably a polynucleotide consisting of the base sequence shown in any one of SEQ ID NOs: 30 to 58, and more preferably a polynucleotide consisting of the base sequence shown in any one of SEQ ID NOs: 30 to 45. The nucleotide sequences of SEQ ID NOs: 30 to 45 (gene sequences of Inocle-α InoC), 46 to 51 (gene sequences of Inocle-β InoC), 52 to 55 (gene sequences of Inocle-γ InoC), and 56 to 58 (gene sequences of Inocle-δ InoC) encode the amino acid sequences of SEQ ID NOs: 1 to 16 (amino acid sequences of Inocle-α InoC), 17 to 22 (amino acid sequences of Inocle-β InoC), 23 to 26 (amino acid sequences of Inocle-γ InoC), and 27 to 29 (amino acid sequences of Inocle-δ InoC), respectively.
[0061] A polynucleotide encoding protein (I) can be used as a diagnostic marker for head and neck cancer or colon cancer, similar to protein (I).
[0062] As described later in the Examples, the nucleotide sequences of SEQ ID NOs: 30 to 45 (gene sequences of Inocle-α InoC) were present in lower amounts in patients with head and neck cancer or colorectal cancer than in healthy individuals. Therefore, if the amount of Inocle-α InoC present in a subject's saliva is lower than the reference value, it indicates that the subject is likely to be suffering from head and neck cancer or colorectal cancer. In other words, the amount of nucleotides present consisting of at least one nucleotide sequence selected from the group consisting of the nucleotide sequences represented by any of SEQ ID NOs: 30 to 45 (nucleotide sequences of Inocle-α InoC) can be said to be a marker indicating that the subject is likely to be suffering from head and neck cancer or colorectal cancer.
[0063] Furthermore, InoC is a sequence that characterizes Inocles, and since InoC is highly identical between Inocle-α, Inocle-β, Inocle-γ, and Inocle-δ, which are types of Inocles, it is understood that InoC of Inocle-α, Inocle-β, Inocle-γ, and Inocle-δ can also be used as markers for head and neck cancer or colorectal cancer. The nucleotide sequences of SEQ ID NOs: 46 to 51 (gene sequence of Inocle-β InoC), SEQ ID NOs: 52 to 55 (gene sequence of Inocle-γ InoC), and SEQ ID NOs: 56 to 58 (gene sequence of Inocle-δ InoC) encode Inocle-β InoC, Inocle-γ InoC, and Inocle-δ InoC proteins that are highly identical to Inocle-α InoC protein. Therefore, it is understood that the amount of nucleotides consisting of at least one nucleotide sequence selected from the group consisting of the nucleotide sequences represented by any of SEQ ID NOs: 30 to 58 can be used as a marker for head and neck cancer or colorectal cancer.
[0064] <Circular DNA> The present specification discloses a circular DNA (hereinafter also referred to as circular DNA (I)) containing polynucleotide (I). The circular DNA (I) of the present invention is not limited to any other base sequence as long as it contains a polynucleotide encoding the protein (I).
[0065] The total length of the circular DNA (I) of the present invention is not particularly limited, but is preferably 100 Kbp to 500 Kbp, and may be, for example, 200 Kbp to 400 Kbp, or 250 Kbp to 350 Kbp.
[0066] The circular DNA (I) is preferably an intracellular ECE (extrachromosomal gene) of an oral bacterium, more preferably an intracellular ECE of Streptococcus (streptococcus, Lactobacillales order, genus Streptococcus), and even more preferably an intracellular ECE of Streptococcus salivarius.
[0067] The circular DNA (I) preferably contains, as another base sequence, a polynucleotide encoding one or more proteins selected from a group of proteins consisting of the amino acid sequences shown in SEQ ID NOs: 61 to 148 or amino acid sequences having 50% or more identity to any of the amino acid sequences shown in SEQ ID NOs: 61 to 148. As shown in Table 1, the amino acid sequences represented by SEQ ID NOs: 61 to 148 are representative amino acid sequences deduced from the base sequences of 38 functional genes (functional gene Nos. 1 to 38) other than Inoc of Inocles.
[0068] [Table 1-1]
[0069] [Table 1-2]
[0070] Functional genes Nos. 1 to 38 share 50% or more identity in terms of amino acid sequence among the 29 types of inocles. Therefore, circular DNA (I) may contain, as other base sequences, polynucleotides encoding one or more proteins selected from the group of proteins consisting of the amino acid sequences represented by SEQ ID NOs: 61 to 148, or may contain polynucleotides encoding one or more proteins selected from the group of proteins consisting of amino acid sequences having 50% or more identity with any of the amino acid sequences represented by SEQ ID NOs: 61 to 148.
[0071] More preferably, the circular DNA (I) contains, as other base sequences, all of the polynucleotides encoding one or more proteins selected from the group of proteins consisting of Inocles functional gene Nos. 1 to 10 shown in Table 1, i.e., amino acid sequences represented by SEQ ID NOs: 61 to 89, or amino acid sequences having 50% or more identity to any of the amino acid sequences represented by SEQ ID NOs: 61 to 89, and further contains at least one polynucleotide encoding one or more proteins selected from the group of proteins consisting of Inocles functional gene (other than Inoc) Nos. 11 to 20, i.e., amino acid sequences represented by SEQ ID NOs: 90 to 112, or amino acid sequences having 50% or more identity to any of the amino acid sequences represented by SEQ ID NOs: 90 to 112.
[0072] More preferably, the circular DNA (I) contains, as other base sequences, all of the polynucleotides encoding one or more proteins selected from the group of proteins consisting of the Inocles functional genes Nos. 21 and 22 shown in Table 1, i.e., the amino acid sequences shown in SEQ ID NOs: 113 to 125, or amino acid sequences having 50% or more identity to any of the amino acid sequences shown in SEQ ID NOs: 113 to 125, and further contains at least one polynucleotide encoding one or more proteins selected from the group of proteins consisting of the Inocles functional genes (other than Inoc) Nos. 23 to 30, i.e., the amino acid sequences shown in SEQ ID NOs: 126 to 140, or amino acid sequences having 50% or more identity to any of the amino acid sequences shown in SEQ ID NOs: 126 to 140.
[0073] More preferably, the circular DNA (I) contains, as other base sequences, all of the polynucleotides encoding one or more proteins selected from the group of proteins consisting of Inocles functional genes Nos. 31 and 32 shown in Table 1, i.e., the amino acid sequences shown in SEQ ID NOs: 141 to 142, or amino acid sequences having 50% or more identity to any of the amino acid sequences shown in SEQ ID NOs: 141 to 142, and further contains at least one polynucleotide encoding one or more proteins selected from the group of proteins consisting of Inocles functional genes Nos. 33 to 38, i.e., the amino acid sequences shown in SEQ ID NOs: 143 to 148, or amino acid sequences having 50% or more identity to any of the amino acid sequences shown in SEQ ID NOs: 143 to 148.
[0074] More preferably, the circular DNA (I) contains, as other base sequences, all of the functional genes Nos. 1 to 10, 21, 22, 31, and 32 of Inocles shown in Table 1, and further contains at least one selected from Nos. 11 to 20, at least one selected from Nos. 22 to 30, and at least one selected from Nos. 33 to 38. That is, the polynucleotides include all polynucleotides encoding one or more proteins selected from the group of proteins consisting of the amino acid sequences shown in SEQ ID NOs: 61 to 89, 113 to 125, and 141 to 142, or amino acid sequences having 50% or more identity to any of the amino acid sequences shown in SEQ ID NOs: 61 to 89, 113 to 125, and 141 to 142, and further include at least one polynucleotide encoding one or more proteins selected from the group of proteins consisting of the amino acid sequences shown in SEQ ID NOs: 90 to 112, or amino acid sequences having 50% or more identity to any of the amino acid sequences shown in SEQ ID NOs: 90 to 112, at least one polynucleotide encoding one or more proteins selected from the group of proteins consisting of the amino acid sequences shown in SEQ ID NOs: 126 to 140, or amino acid sequences having 50% or more identity to any of the amino acid sequences shown in SEQ ID NOs: 126 to 140, and at least one polynucleotide encoding one or more proteins selected from the group of proteins consisting of the amino acid sequences shown in SEQ ID NOs: 143 to 144, or amino acid sequences having 50% or more identity to any of the amino acid sequences shown in SEQ ID NOs: 143 to 144.
[0075] Preferably, the circular DNA (I) further comprises, as other base sequences, at least one selected from the group consisting of polynucleotides encoding VirB4 or VirD4. Preferably, the circular DNA (I) further comprises, as another base sequence, at least one selected from the group consisting of polynucleotides encoding LPXTG-like motif PF. The LPXTG-like motif PF is preferably an amino acid sequence represented by SEQ ID NOs: 180 to 185.
[0076] The circular DNA (I) preferably contains genes homologous to Streptococcus chromosomal genes, more preferably 40 or more, more preferably 45 or more, and even more preferably 50 or more of the homologous genes. The homologous genes can be identified by performing a similarity search between the ORFs of the circular DNA (I) and the ORFs of the Streptococcus chromosome using DIAMOND (v2.1.8) with the options --sensitive and e-value<1e-10. As the gene homologous to a chromosomal gene of Streptococcus, a RecD-like DNA helicase gene, a DNA polymerase III subunit α gene, a bifunctional transglycosylase gene, or a sortase A gene is preferred.
[0077] The base sequence of the circular DNA (I) is preferably any of the base sequences shown in SEQ ID NOs: 149 to 177. The base sequences shown in SEQ ID NOs: 149 to 177 are the base sequences of Inocles 001 to 0029, respectively.
[0078] The circular DNA (I), like the protein (I) and the polynucleotide (I), can be used as a diagnostic marker for head and neck cancer or colon cancer.
[0079] <Cancer screening methods> A method for testing for head and neck cancer or colorectal cancer, which includes a step (C1) of measuring the amount of Inocles present in the saliva of a subject, is referred to as method (C). According to method (C), it is possible to determine whether a subject is suffering from head and neck cancer or colorectal cancer in a simple and non-invasive manner.
[0080] Head and neck cancer is cancer that occurs in the head and neck region, and includes lip and oral cavity cancer (tongue cancer, floor of the mouth cancer, gum cancer, etc.), laryngeal cancer, pharyngeal cancer (nasopharyngeal cancer, oropharyngeal cancer, hypopharyngeal cancer), nasal and paranasal sinus cancer (nasal cavity cancer, maxillary sinus cancer, etc.), salivary gland cancer (parotid gland cancer, submandibular gland cancer, etc.), and thyroid cancer.
[0081] Colorectal cancer is cancer that occurs in the colon (cecum, ascending colon, transverse colon, descending colon, sigmoid colon) or rectum (upper rectum, lower rectum), and includes adenocarcinoma, squamous cell carcinoma, and adenosquamous carcinoma. The head and neck cancer or colorectal cancer in the present invention may be at any of the following clinical stages: 0, IA1, IA2, IA3, IB, IIA, IIB, IIIA, IIIB, IIIC, IVA, and IVB.
[0082] As used herein, the term "testing method" refers to examining a sample collected from a subject to obtain information necessary for diagnosis, and the testing method of the present invention can be carried out, for example, by a testing company. As used herein, the term "method for testing head and neck cancer or colon cancer" refers to examining head and neck cancer or colon cancer and any phenomena caused by them, and means examining a sample collected from a subject to obtain information necessary for determining, for example, the presence or absence of the disease, the stage of progression, the degree of malignancy, whether progression has been stopped or delayed, the presence or absence of metastasis, and the presence or absence of treatment effect. In this specification, the "method for testing head and neck cancer or colon cancer" may be used for so-called companion diagnosis, which predicts the effects and side effects of anticancer drugs.
[0083] The head and neck cancer or colon cancer testing method of the present invention can be rephrased as "a head and neck cancer or colon cancer testing method that provides data related to head and neck cancer or colon cancer," "a method for determining the possibility that a subject has head and neck cancer or colon cancer," "a method for evaluating the possibility that a subject has head and neck cancer or colon cancer," "a method for testing the possibility that a subject has head and neck cancer or colon cancer," "a method for determining whether a subject has head and neck cancer or colon cancer," or "a method for determining head and neck cancer or colon cancer." Furthermore, the head and neck cancer or colon cancer testing method of the present invention may provide test results that serve as material for a physician to diagnose head and neck cancer in a subject, and can also be rephrased as "a method for collecting data for diagnosing head and neck cancer or colon cancer," "a method for analyzing diagnostic information to infer that a subject has head and neck cancer or colon cancer," or "a method for analyzing diagnostic information to assist in inferring that a subject has head and neck cancer or colon cancer." The present specification also discloses a method for diagnosing head and neck cancer or colorectal cancer. The method for diagnosing cancer means that a doctor or other doctor uses the results obtained by the testing method of the present invention to determine the presence or absence of cancer, the stage of progression, the malignancy, whether progression has been stopped or delayed, the presence or absence of metastasis, and whether treatment has been effective.
[0084] <Process (C1)> Step (C1) is a step of measuring the amount of Inocles present in the saliva of a subject.
[0085] The terms "subject," "individual," or "patient" can be used interchangeably herein and generally and preferably refer to humans, but can also include reference to non-human animals, preferably warm-blooded animals, and even more preferably mammals, such as non-human primates, rodents, dogs, cats, horses, sheep, pigs, and the like. The term "non-human animal" includes all vertebrates, e.g., mammals such as non-human primates (especially higher primates), sheep, dogs, rodents (e.g., mice or rats), guinea pigs, goats, pigs, cats, rabbits, and cows, and non-mammals such as chickens, amphibians, and reptiles. In certain embodiments, the subject is a non-human mammal. In certain embodiments, the subject is a human subject. The term does not denote a particular age or sex. Thus, adult and newborn subjects, as well as fetuses, regardless of male or female, are intended to be encompassed. Examples of subjects include humans, dogs, cats, cows, goats, and mice.
[0086] Saliva can be collected from a subject by a known method and then processed by a method suitable for detecting the amount of Inocles present before use. Such a processing method preferably includes the method for producing a nucleic acid solution (X) described below.
[0087] The Inocles to be detected may be any of Inocle-α, Inocle-β, Inocle-γ, and Inocle-δ, but are preferably Inocle-α. The origin of the detected Inocles is not particularly limited, but is preferably derived from oral bacteria, more preferably from Streptococcus, and even more preferably from Streptococcus salivarius.
[0088] The method for measuring the abundance of Inocles in step (C1) is not particularly limited, and any known method can be used. Examples of such methods include a method using a next-generation sequencer (NGS), a method using primers capable of specifically amplifying Inocles, and a method using a probe that specifically hybridizes to Inocles DNA. Examples of Inocles-specific primers or probes include the primer set (I) and probe (I) described below.
[0089] Since Inocles comprise a very large number of bases, it may be difficult to confirm the abundance of the entire base sequence. Therefore, step (C1) is preferably a step of measuring the abundance of a portion of the base sequence of Inocles, more preferably a step of measuring the abundance of InoC.
[0090] The abundance of InoC may preferably be the abundance of one or more polynucleotides selected from the group of polynucleotides consisting of the base sequences represented by SEQ ID NOs: 30 to 58 (base sequences of InoC of Inocle-α, β, γ, and δ), or may be the total abundance of the group of polynucleotides consisting of the base sequences represented by SEQ ID NOs: 30 to 58, but is preferably the abundance of one or more polynucleotides selected from the group of polynucleotides consisting of the base sequences represented by SEQ ID NOs: 30 to 45 (base sequences of Inocle-α InoC), and more preferably the total abundance of the group of polynucleotides consisting of the base sequences represented by SEQ ID NOs: 30 to 45.
[0091] When measuring the abundance of Inocles using a next-generation sequencer, the value obtained by normalizing the number of Inocle-related reads by the total number of reads in each sample can be used. This value may be determined based on RPKM (reads per kilobase of contig per million total reads), RPM (Reads Per Million mapped reads), FPKM (Fragments Per Kilobase of exon per Million mapped reads), etc., but is preferably determined based on RPKM, and more preferably is the total RPKM of InoCs in Inocle-α.
[0092] In the method (C), preferably, the abundance of Inocles obtained in the step (C1) being smaller than the reference value indicates that the subject is likely to be affected with head and neck cancer or colorectal cancer. As will be described later in the Examples, the amount of Inocles present in patients with head and neck cancer or colorectal cancer is lower than that in healthy individuals. Therefore, if the amount of Inocles present in the step (C1) is lower than the reference value, it can be inferred, determined, or diagnosed that the subject is highly likely to be suffering from head and neck cancer or colorectal cancer.
[0093] <Other processes> The method (C) preferably further comprises the step (C4) of calculating a reference value. Step (C4) preferably includes a step (C4-1) of measuring the amount of Inocles present in the saliva of a healthy subject, and a step (C4-2) of calculating a reference value from the amount present obtained in step (C4-1). The method (C) preferably further comprises a step (C5) of comparing the abundance of Inocles obtained in the step (C1) with the reference value obtained in the step (C4-2). Step (C4-1) can be carried out in the same manner as step (C1) using saliva from a healthy subject.
[0094] The reference value is preferably the abundance of Inocles obtained from healthy subjects. The reference value is preferably a value calculated from multiple measurement results obtained from a group of multiple healthy subjects, or a numerical range, such as the mean, median, or mean ± standard deviation. A healthy individual refers to a person who is not suspected of having head and neck cancer or colorectal cancer, or who has not been diagnosed with head and neck cancer or colorectal cancer.
[0095] <Method (Ci)> Method (C) may further comprise a step (C2) of measuring the amount of Streptcoccus Infantis present in the saliva of the subject, and a step (C3) of calculating the amount of Inocles present per Streptcoccus Infantis by dividing the amount of Inocles present obtained in step (C1) by the amount of Streptcoccus Infantis present in step (C2). In this case, method (C) is also referred to as method (Ci).
[0096] Step (C2) is a step of measuring the amount of Streptococcus Infantis present in the saliva of the subject. The method for measuring the abundance of Streptcoccus Infantis is not particularly limited, and may be performed by a known method, such as a method using a next-generation sequencer (NGS), a method using a marker that specifically recognizes Streptcoccus Infantis, or a method that counts the number of Streptcoccus Infantis.
[0097] When measuring the abundance of Streptcoccus Infantis using a next-generation sequencer, a value obtained by normalizing the number of reads related to Streptcoccus Infantis by the total number of reads in each sample can be used. This value may be determined based on RPKM (reads per kilobase of contig per million total reads), RPM (reads per million mapped reads), FPKM (fragments per kilobase of exon per million mapped reads), etc., but is preferably determined based on RPKM, and more preferably the total RPKM of Streptcoccus Infantis.
[0098] Step (C3) is a step of calculating the amount of Inocles present per Streptcoccus Infantis by dividing the amount of Inocles present obtained in step (C1) by the amount of Streptcoccus Infantis present obtained in step (C2). In step (C3), the calculation is preferably carried out by dividing the RPKM value of Inocles by the RPKM value of Streptcoccus Infantis, more preferably by dividing the RPKM value of Inocle-α by the RPKM value of Streptcoccus Infantis.
[0099] In the method (Ci), preferably, the abundance of Inocles per Streptococcus Infantis obtained in the step (C3) being smaller than the reference value indicates that the subject is likely to be suffering from head and neck cancer or colorectal cancer. As will be described later in the Examples, the amount of Inocles present per Streptcoccus Infantis in patients with head and neck cancer or colon cancer is lower than that in healthy individuals. Therefore, if the amount of Inocles present per Streptcoccus Infantis obtained in step (C3) is lower than the reference value, it can be inferred, determined, or diagnosed that the subject is highly likely to be suffering from head and neck cancer or colon cancer.
[0100] In method (Ci), the step (C4) of calculating the reference value preferably includes a step (C4i-1) of measuring the abundance of Inocles contained in the saliva of healthy individuals, a step (C4i-2) of measuring the abundance of Streptcoccus Infantis contained in the saliva of healthy individuals, a step (C4i-3) of dividing the abundance of Inocles obtained in the step (C4i-1) by the abundance of Streptcoccus Infantis obtained in the step (C4i-2) to calculate the abundance of Inocles per Streptcoccus Infantis, and a step (C4i-4) of calculating the reference value from the abundance of Inocles per Streptcoccus Infantis obtained in the step (C4i-3). Method (Ci) preferably further includes a step (Ci5) of comparing the abundance of Inocles per Streptcoccus Infantis obtained in the step (C3) with the reference value obtained in the step (C4i-4). Step (C4i-1) can be performed in the same manner as step (C1), step (C4i-2) can be performed in the same manner as step (C2), and step (C4i-3) can be performed in the same manner as step (C3).
[0101] In method (Ci), the reference value is preferably the abundance of Inocles per Streptcoccus Infantis obtained from a subject who is a healthy individual. The reference value is preferably a value or numerical range calculated from a plurality of measurement results obtained from a plurality of groups of healthy individuals, and examples include an average value, a median value, an average value ± standard deviation, etc.
[0102] <Saliva test kit for head and neck cancer or colorectal cancer detection> This specification discloses a saliva test kit for head and neck cancer or colorectal cancer detection (hereinafter, also referred to as saliva test kit (I)) containing an Inocles measurement reagent.
[0103] <Inocles measurement reagent> The Inocles measurement reagent is a reagent for measuring the amount of Inocles present in saliva, and is not particularly limited as long as it can measure the amount of Inocles present in saliva. Details regarding the measurement of inocles abundance are the same as those in the section on "Testing methods for head and neck cancer or colorectal cancer."
[0104] Examples of the reagent for measuring Inocles include Inocles-specific primers or probes, etc. Examples of the Inocles-specific primers or probes include the primer set (I) or probe (I) described below.
[0105] The Inocles measurement reagent may further contain known optional components such as stabilizers, pH adjusters, and emulsifiers, as long as they do not impair the effects of the present invention.
[0106] A saliva test kit for head and neck cancer or colorectal cancer testing, which includes an Inocles measurement reagent, may include, in addition to the Inocles measurement reagent, other reagents for washing, concentration, dilution, fixation, drying, preservation, staining, color development, observation, etc., and may also include laboratory equipment, etc. as needed. The laboratory equipment may be an equipment for collecting saliva, such as a swab. A saliva test kit for head and neck cancer or colon cancer testing, which contains an Inocles measurement reagent, preferably further contains a nuclease. The saliva test kit (I) may further contain, in addition to the Inocles measurement reagent, one or more reagents selected from the group consisting of DNA polymerase, nucleotides, and nucleic acid amplification buffers.
[0107] The saliva test kit (I) may be used by the subject to be tested, or by a medical professional (doctor, nurse, pharmacist, clinical laboratory technician, etc.).
[0108] Because the saliva test kit (I) is simple and inexpensive, it is preferably used for colonoscopy, nasopharyngeal endoscopy, CT scan, panoramic X-ray, MRI scan, or pre-stage testing before biopsy. It is more preferable to use the saliva test kit (I) to recommend a medical interview by a specialist physician, colonoscopy, nasopharyngeal endoscopy, CT scan, panoramic X-ray, MRI scan, or biopsy when the saliva test kit (I) indicates that a subject may be suffering from head and neck cancer or colorectal cancer.
[0109] The saliva test kit (I) may further contain a reagent for measuring Streptococcus Infantis. The Streptcoccus Infantis measurement reagent is a reagent for measuring the amount of Streptcoccus Infantis present in saliva, and is not particularly limited as long as it can measure the amount of Streptcoccus Infantis present in saliva. Details regarding the measurement of the abundance of Streptococcus Infantis are the same as those in the section on "Testing methods for head and neck cancer or colorectal cancer."
[0110] Examples of the reagent for measuring Streptcoccus Infantis include primers or probes specific to Streptcoccus Infantis, which may be primers or probes specific to Streptcoccus Infantis used in quantitative RT-PCR, quantitative (real-time) PCR, microRNA array, RNA sequencing (RNA-Seq), multiplex miRNA profiling, or the like. The Streptococcus Infantis measurement reagent may further contain known optional components such as stabilizers, pH adjusters, and emulsifiers, as long as they do not impair the effects of the present invention.
[0111] <marker> The present specification discloses diagnostic markers for head and neck cancer or colorectal cancer, which comprise proteins (I), polynucleotides (I), or circular DNAs (I). The present specification discloses markers for determining whether a patient is affected with head and neck cancer or colorectal cancer, the markers comprising proteins (I), polynucleotides (I), or circular DNAs (I). The present specification discloses a protein (I), a polynucleotide (I), or a circular DNA (I) that can be used as a diagnostic marker for head and neck cancer or colorectal cancer.
[0112] As described later in the Examples, the amount of Inocle-α present in the saliva of patients with head and neck cancer or colorectal cancer is lower than that of healthy individuals. Therefore, Inocles can be used as a marker for determining head and neck cancer or colorectal cancer, preferably as a marker for determining and diagnosing whether or not a patient has head and neck cancer or colorectal cancer.
[0113] <Production method of nucleic acid solution (X)> This specification discloses a method for producing a nucleic acid solution derived from oral bacteria (hereinafter also referred to as method (X) for producing a nucleic acid solution), which comprises a step (A1) of centrifuging an animal-derived specimen containing said oral bacteria, and a step (A2) of treating with a nuclease the pellet obtained after removing the supernatant from the specimen after step (A1). According to the method (X) for producing a nucleic acid solution, contamination with animal-derived nucleic acids is reduced, and a nucleic acid solution containing nucleic acids derived from the oral bacteria at a high purity can be produced.
[0114] <Process (A1)> Step (A1) is a step of centrifuging an animal-derived specimen containing oral bacteria. Oral bacteria are not particularly limited as long as they are bacteria present in the oral cavity, and include, for example, bacteria of the genus Streptococcus such as Streptococcus salivarius and Streptococcus Infantis.
[0115] The animal is preferably a human, but may also include reference to a non-human animal, preferably a warm-blooded animal, and even more preferably a mammal, such as a non-human primate, rodent, dog, cat, horse, sheep, pig, etc. The term "non-human animal" includes all vertebrates, for example mammals such as non-human primates (especially higher primates), sheep, dog, rodent (e.g., mouse or rat), guinea pig, goat, pig, cat, rabbit, cow, and non-mammals such as chickens, amphibians, reptiles, etc.
[0116] The animal-derived specimen containing oral bacteria is not particularly limited, and examples thereof include saliva, oral swabs, etc., with saliva being preferred. The method of centrifugation is not particularly limited, and any known method may be used. The conditions for centrifugation are not particularly limited, but centrifugation is preferably carried out at 5,000×g or more, more preferably 10,000×g or more. The time for which centrifugation is carried out is not particularly limited, but is preferably 1 minute or more, more preferably 5 minutes or more, and even more preferably 10 minutes or more.
[0117] If necessary, animal-derived specimens containing oral bacteria may be diluted before centrifugation. The diluent used for dilution is not particularly limited as long as it does not damage the oral bacteria, and examples of the diluent include water and buffer solutions. Examples of the buffer solution include buffer solutions commonly used in biochemical tests, such as Tris buffer solution and phosphate buffer solution.
[0118] <Process (A2)> Step (A2) is a step of treating the pellet obtained after removing the supernatant from the sample after step (A1) with a nuclease. The method for removing the supernatant is not particularly limited and can be carried out by a known method, such as decantation or removal with a pipette.
[0119] Nuclease treatment may be carried out by a known method, and the type of nuclease used and the treatment conditions are not particularly limited. The temperature for nuclease treatment is preferably room temperature or higher, more preferably around 37°C, and the treatment time is preferably 10 minutes or longer, more preferably 20 minutes or longer, and even more preferably 30 minutes or longer. The amount or concentration of the nuclease is not particularly limited as long as it can sufficiently decompose the nucleic acids contained in the pellet portion, and can be changed appropriately depending on the type of nuclease used and the type and amount of the sample.
[0120] The type of nuclease is not particularly limited, and may be an exonuclease, an endonuclease, or a combination thereof. Examples of nucleases include DNase and RNase, and preferably both DNase and RNase are used. The DNase is preferably DNaseI, and the RNase is preferably RNaseI.
[0121] The pellet obtained after removing the supernatant from the sample after step (A1) may be dispersed in a solvent before being treated with a nuclease. The dispersion liquid used for dispersion is not particularly limited as long as it does not damage oral bacteria and does not inhibit nuclease activity, and examples of the dispersion liquid include water and buffer solutions. As the buffer solution, buffer solutions commonly used in biochemical tests can be used, such as Tris buffer solution and phosphate buffer solution. The dispersion liquid may contain Mg to improve nuclease activity. 2+ It is preferred that the compound contains:
[0122] The production method (X) preferably comprises, after the step (A2), a step of extracting and purifying nucleic acid from the sample after the step (A2). The method for extracting and purifying nucleic acids is not particularly limited, and known methods can be used. Specifically, nucleic acids can be extracted and purified by the methods described in the Examples below.
[0123] The production method (X) may include a step of centrifuging the nuclease-treated sample after step (A2) and before the nucleic acid extraction and purification step, which can increase the concentration of oral bacteria in the sample to be subjected to the nucleic acid extraction and purification step.
[0124] The production method (X) may include a step of treating the sample obtained after step (A2) with a lytic enzyme after step (A2) but before the step of extracting and purifying nucleic acids. Examples of the lytic enzyme include lysozyme and achromopeptidase. This step lyses oral bacteria in the sample, allowing for efficient extraction and purification of nucleic acids from the oral bacteria.
[0125] The production method (X) may include a step of treating the sample obtained after step (A2) with a protease, after step (A2) and before the step of extracting and purifying nucleic acids. Examples of the protease include Proteinase K. This step allows for rapid inactivation of nucleases and lytic enzymes.
[0126] Animal-derived samples containing oral bacteria (e.g., human saliva) contain large amounts of animal-derived nucleic acids in addition to nucleic acids derived from oral bacteria, and 85-95% of the DNA contained in the sample is genomic DNA of animal origin, making it difficult to obtain sufficient sequencing depth for oral bacteria from such samples. However, according to production method (X), a nucleic acid solution derived from oral bacteria, in which the content of animal-derived nucleic acids is reduced, can be easily produced from an animal-derived specimen containing oral bacteria. This nucleic acid solution also minimizes the impact on the oral bacterial composition. Because the nucleic acid solution produced by production method (X) has the advantages described above, it is easy to obtain sufficient sequencing depth for oral bacteria.
[0127] <Primer set> The present specification discloses a primer set (hereinafter also referred to as primer set (I)) that specifically recognizes and amplifies at least a part of a polynucleotide consisting of the nucleotide sequence shown in SEQ ID NOs: 149 to 177. Primer set (I) can be used to detect Inocles.
[0128] The sequence and length of primer set (I) are not limited as long as it hybridizes with at least a portion of a polynucleotide consisting of the base sequence represented by SEQ ID NOs: 149 to 177 and can amplify a portion of the base sequence represented by SEQ ID NOs: 149 to 177. The primer set (I) usually consists of a forward primer and a reverse primer. Each primer of the primer set (I) can be synthesized using a method known in the art.
[0129] Primer set (I) can hybridize with Inocles in a sample to amplify at least a portion of the base sequence of the Inocles using the Inocles in the specimen as a template, and can also detect Inocles by hybridization using the primer set as a probe.
[0130] Each primer in primer set (I) may have an additional base sequence or may be labeled, as long as it does not lose its specificity for at least a portion of the polynucleotide consisting of the base sequences represented by SEQ ID NOs: 149 to 177. The label is not particularly limited as long as it is a substance that can normally be used to label nucleic acids, and examples thereof include fluorescent substances, low-molecular-weight compounds such as biotin, and radioisotopes. Examples of the additional base sequence include restriction enzyme recognition sequences and various promoter sequences.
[0131] Primer set (I) may be a primer set (also referred to as primer set (Ia)) that specifically recognizes and amplifies InoC, i.e., at least a portion of a polynucleotide consisting of a base sequence represented by any one of SEQ ID NOs: 30 to 58, or may be a primer set (also referred to as primer set (Ib)) that specifically recognizes and amplifies at least a portion of InoC, i.e., at least a portion of a polynucleotide consisting of a base sequence represented by any one of SEQ ID NOs: 149 to 177, other than the base sequence represented by any one of SEQ ID NOs: 30 to 58.
[0132] The primer set (Ia) is preferably selected from the following primer sets (Ia-A) to (Ia-D). Primer set (Ia-A): a polynucleotide consisting of the nucleotide sequence represented by SEQ ID NO: 200 and a polynucleotide consisting of the nucleotide sequence represented by SEQ ID NO: 201 Primer set (Ia-B): a polynucleotide consisting of the nucleotide sequence represented by SEQ ID NO: 202 and a polynucleotide consisting of the nucleotide sequence represented by SEQ ID NO: 203 Primer set (Ia-C): a polynucleotide consisting of the nucleotide sequence represented by SEQ ID NO: 205 and a polynucleotide consisting of the nucleotide sequence represented by SEQ ID NO: 206 Primer set (Ia-D): a polynucleotide consisting of the nucleotide sequence represented by SEQ ID NO: 208 and a polynucleotide consisting of the nucleotide sequence represented by SEQ ID NO: 209
[0133] The primer set (Ib) is preferably selected from the following primer sets (Ib-A) to (Ib-E). Primer set (Ib-A): a polynucleotide consisting of the nucleotide sequence represented by SEQ ID NO: 190 and a polynucleotide consisting of the nucleotide sequence represented by SEQ ID NO: 191 Primer set (Ib-B): a polynucleotide consisting of the nucleotide sequence represented by SEQ ID NO: 192 and a polynucleotide consisting of the nucleotide sequence represented by SEQ ID NO: 193 Primer set (Ib-C): a polynucleotide consisting of the nucleotide sequence represented by SEQ ID NO: 194 and a polynucleotide consisting of the nucleotide sequence represented by SEQ ID NO: 195 Primer set (Ib-D): a polynucleotide consisting of the nucleotide sequence represented by SEQ ID NO: 196 and a polynucleotide consisting of the nucleotide sequence represented by SEQ ID NO: 197 Primer set (Ib-E): a polynucleotide consisting of the nucleotide sequence shown in SEQ ID NO: 198 and a polynucleotide consisting of the nucleotide sequence shown in SEQ ID NO: 199.
[0134] <probe> The present specification discloses a probe (hereinafter also referred to as probe (I)) capable of specifically binding to at least a part of a polynucleotide consisting of the base sequence shown in SEQ ID NOs: 149 to 177. Probe (I) is used to detect Inocles. The sequence and length of probe (I) are not limited as long as it hybridizes with at least a portion of a polynucleotide consisting of the base sequence represented by SEQ ID NOs: 149 to 177 and can detect a portion of the base sequence represented by SEQ ID NOs: 149 to 177. The probe (I) can be synthesized using techniques known in the art.
[0135] The probe (I) can capture the Inocles in the specimen by hybridizing the probe (I) with the Inocles in the sample. The Inocles captured using the probe (I) can be quantified, for example, by a known real-time detection PCR method.
[0136] Probe (I) may have an additional base sequence and may be labeled, as long as it does not lose its specificity for at least a part of the polynucleotide consisting of the base sequences represented by SEQ ID NOs: 149 to 177. The label is not particularly limited as long as it is a substance that can normally be used to label nucleic acids, and examples of the label include fluorescent substances, low-molecular-weight compounds such as biotin, and radioisotopes. Further, the probe (I) may be immobilized on the surface of, for example, a probe-immobilized solid support (e.g., labeled particles, magnetic particles, etc.).
[0137] The probe (I) may be a probe (also referred to as probe (Ia)) that can specifically bind to at least a part of a polynucleotide consisting of a base sequence represented by any one of SEQ ID NOs: 30 to 58, i.e., InoC, or a probe (also referred to as probe (Ib)) that can specifically bind to at least a part of a portion other than InoC of Inocles, i.e., a portion other than the base sequence represented by any one of SEQ ID NOs: 30 to 58 among the polynucleotides consisting of the base sequence represented by SEQ ID NOs: 149 to 177.
[0138] The probe (Ia) is preferably selected from the following probes (Ia-A) to (Ia-C). Probe (Ia-A): A polynucleotide consisting of the base sequence represented by SEQ ID NO: 204 Probe (Ia-B): A polynucleotide consisting of the base sequence represented by SEQ ID NO: 207 Probe (Ia-C): A polynucleotide consisting of the base sequence represented by SEQ ID NO: 210
[0139] <Method for detecting Inocles> This specification discloses a method for detecting Inocles (hereinafter also referred to as detection method (I)), which includes a step of hybridizing a primer set (I) or a probe (I) with DNA contained in saliva (hereinafter also referred to as step (D)). For step (D), as long as the primer set (I) or the probe (I) can be hybridized with the DNA contained in saliva, the method is not limited. The hybridization conditions can be appropriately selected in consideration of the number of amplified bases of the nucleic acid, Tm value, base sequence, etc.
[0140] When using the primer set (I), Inocles can be detected by nucleic acid amplification using DNA contained in saliva as a template, or by subjecting the DNA contained in saliva to hybridization using the primer set as a probe. Examples of the detection method by nucleic acid amplification include a method of determining the nucleotide sequence of the nucleic acid amplification product by electrophoresis, a sequencer, or the like. Examples of the detection method by hybridization include spot hybridization, colony hybridization, in situ hybridization, nucleic acid sandwich hybridization, and the like.
[0141] When using the probe (I), Inocles can be detected by hybridizing with the DNA contained in saliva. More specifically, examples of the detection method include spot hybridization, colony hybridization, in situ hybridization, nucleic acid sandwich hybridization, and the like.
[0142] <Kit for detecting Inocles> This specification discloses a kit for detecting Inocles (hereinafter referred to as kit (I)) containing the primer set (I) and / or the probe (I). In addition to the primer set (I) and / or the probe (I), the kit (I) may further contain one or more reagents selected from the group consisting of DNA polymerase, nucleotides, and a buffer for nucleic acid amplification. The DNA polymerase can be selected and used by appropriately considering the number of amplified bases of the nucleic acid, the Tm value, the nucleotide sequence, and the like. The nucleotides may be deoxynucleotides labeled with a labeling reagent or the like. The buffer for nucleic acid amplification can be appropriately adjusted and used from known ones in consideration of the DNA polymerase to be used.
[0143] This specification discloses the inventions relating to the following
[01] to
[503] . Note that the term "Inocle-related molecules" is used to mean Inocles or parts of Inocles, or proteins encoded by them.
[01] A step (A1) of centrifuging an animal-derived specimen containing oral bacteria to remove the supernatant and obtain a pellet portion; a step (A2) of treating the pellet obtained in the step (A1) with a nuclease; A method for producing a nucleic acid solution derived from oral bacteria, comprising:
[02] The manufacturing method described in
[01] , wherein the sample is saliva.
[03] The method according to
[01] or
[02] , wherein the animal is a human.
[0144]
[0101] A protein consisting of any one of the following amino acid sequences (a) to (c): (a) an amino acid sequence represented by any one of SEQ ID NOs: 1 to 29 (b) an amino acid sequence having an identity of 90% or more but less than 100% with an amino acid sequence represented by any one of SEQ ID NOs: 1 to 29 (c) an amino acid sequence represented by any one of SEQ ID NOs: 1 to 29, in which 1 to 10 amino acids are deleted, substituted, inserted, or added;
[0102]
[0101] A polynucleotide encoding the protein described in
[0103] A circular DNA comprising the polynucleotide described in
[0101] .
[0104] The circular DNA according to
[0103] , having a total length of 100 Kbp to 500 Kbp.
[0145]
[0201] A cancer testing method comprising a step (B1) of measuring the expression level of Inocle-related molecules contained in the saliva of a subject.
[0202] The examination method described in
[0201] , wherein the cancer is head and neck cancer or colorectal cancer.
[0203] Furthermore, a step (B2) of measuring the amount of Streptococcus Infantis contained in the saliva of the subject; and a step (B3) of calculating the expression level of the Inocle-related molecule per Streptococcus Infantis by dividing the expression level of the Inocle-related molecule obtained in the step (B1) by the amount of Streptococcus Infantis obtained in the step (B2); An inspection method described in
[0201] or
[0202] , which includes:
[0204] The testing method described in any one of
[0201] to
[0203] , wherein the expression level of the Inocle-related molecule obtained in the step (B1) or the expression level of the Inocle-related molecule per Streptococcus Infantis obtained in the step (B3) being smaller than a reference value indicates that the subject is likely to be suffering from cancer.
[0205] The testing method described in
[0204] , wherein the reference value is the expression level of Inocle-related molecules obtained from a healthy subject, or the expression level of Inocle-related molecules per Streptococcus Infantis obtained in step (B3).
[0146]
[0206] The examination method described in any one of
[0201] to
[0205] , wherein the Inocle-related molecule is at least one selected from the group consisting of Inocle-α, Inocle-β, Inocle-γ, and Inocle-δ.
[0207] An examination method described in any one of
[0201] to
[0206] , wherein the Inocle-related molecule is Inocle-α.
[0208] An examination method described in any of
[0201] to
[0207] , wherein the Inocle-related molecule is Inocle-α, and the expression level of the Inocle-related molecule is the expression level of a nucleotide consisting of at least one base sequence selected from the group consisting of base sequences represented by any of SEQ ID NOs: 30 to 45.
[0209] An examination method described in any of
[0201] to
[0208] , wherein the Inocle-related molecule is Inocle-α, and the expression level of the Inocle-related molecule is the sum of the expression levels of all nucleotides consisting of a base sequence represented by any of SEQ ID NOs: 30 to 45.
[0210] The testing method described in any of
[0201] to
[0206] , wherein the Inocle-related molecule is at least one selected from the group consisting of Inocle-α, Inocle-β, Inocle-γ, and Inocle-δ, and the expression level of the Inocle-related molecule is the sum of the expression levels of all nucleotides consisting of a base sequence represented by any of SEQ ID NOs: 30 to 58.
[0147]
[0211] An examination method described in any of
[0201] to
[0206] , wherein the Inocle-related molecule is at least one selected from the group consisting of Inocle-β, Inocle-γ, and Inocle-δ, and the expression level of the Inocle-related molecule is the expression level of a nucleotide consisting of at least one base sequence selected from the group consisting of base sequences represented by any of SEQ ID NOs: 46 to 58.
[0212] An examination method described in any of
[0201] to
[0206] , wherein the Inocle-related molecule is at least one selected from the group consisting of Inocle-β, Inocle-γ, and Inocle-δ, and the expression level of the Inocle-related molecule is the sum of the expression levels of all nucleotides consisting of a base sequence represented by any of SEQ ID NOs: 46 to 58.
[0215] An examination method described in any of
[0201] to
[0212] , wherein the expression level of the Inocle-related molecule is a value determined based on RPKM (reads per kilobase of contig per million total reads).
[0148]
[0216] An examination method described in any one of
[0201] to
[0212] and
[0215] , wherein the amount of Streptococcus Infantis is a value determined based on RPKM (reads per kilobase of contig per million total reads).
[0217] An examination method described in any of
[0201] to
[0207] , wherein the Inocle-related molecule is Inocle-α, and the expression level of the Inocle-related molecule is the expression level of a protein consisting of at least one amino acid sequence selected from the group consisting of amino acid sequences represented by any of SEQ ID NOs: 1 to 16.
[0218] An examination method described in any of
[0201] to
[0207] , wherein the Inocle-related molecule is Inocle-α, and the expression level of the Inocle-related molecule is the total expression level of all proteins consisting of an amino acid sequence represented by any of SEQ ID NOs: 1 to 16.
[0219] An examination method described in any of
[0201] to
[0206] , wherein the Inocle-related molecule is at least one selected from the group consisting of Inocle-α, Inocle-β, Inocle-γ, and Inocle-δ, and the expression level of the Inocle-related molecule is the sum of the expression levels of all proteins consisting of at least one amino acid sequence selected from the group consisting of amino acid sequences represented by any of SEQ ID NOs: 1 to 29.
[0220] An examination method described in any of
[0201] to
[0206] , wherein the Inocle-related molecule is at least one selected from the group consisting of Inocle-β, Inocle-γ, and Inocle-δ, and the expression level of the Inocle-related molecule is the expression level of a protein consisting of at least one amino acid sequence selected from the group consisting of amino acid sequences represented by any of SEQ ID NOs: 17 to 29.
[0149]
[0221] An examination method described in any of
[0201] to
[0206] , wherein the Inocle-related molecule is at least one selected from the group consisting of Inocle-β, Inocle-γ, and Inocle-δ, and the expression level of the Inocle-related molecule is the total expression level of all proteins consisting of an amino acid sequence represented by any of SEQ ID NOs: 17 to 29.
[0150]
[0301] A method for diagnosing cancer, comprising a step (B1) of measuring the expression level of an Inocle-related molecule contained in the saliva of a subject.
[0302] The diagnostic method described in
[0301] , wherein the cancer is head and neck cancer or colon cancer.
[0303] Furthermore, a step (B2) of measuring the amount of Streptococcus Infantis contained in the saliva of the subject; and a step (B3) of calculating the expression level of the Inocle-related molecule per Streptococcus Infantis by dividing the expression level of the Inocle-related molecule obtained in the step (B1) by the amount of Streptococcus Infantis obtained in the step (B2); A diagnostic method described in
[0301] or
[0302] , which includes:
[0304] The diagnostic method described in any one of
[0301] to
[0303] , wherein the expression level of the Inocle-related molecule obtained in the step (B1) or the expression level of the Inocle-related molecule per Streptococcus Infantis obtained in the step (B3) being smaller than a reference value indicates that the subject is highly likely to be suffering from cancer.
[0305] The diagnostic method described in
[0304] , wherein the reference value is the expression level of Inocle-related molecules obtained from a healthy subject, or the expression level of Inocle-related molecules per Streptococcus Infantis obtained in step (B3).
[0151]
[0306] A diagnostic method described in any of
[0301] to
[0305] , wherein the Inocle-related molecule is at least one selected from the group consisting of Inocle-α, Inocle-β, Inocle-γ, and Inocle-δ.
[0307] A diagnostic method described in any of
[0301] to
[0306] , wherein the Inocle-related molecule is Inocle-α.
[0308] A diagnostic method described in any of
[0301] to
[0307] , wherein the Inocle-related molecule is Inocle-α, and the expression level of the Inocle-related molecule is the expression level of a nucleotide consisting of at least one base sequence selected from the group consisting of base sequences represented by any of SEQ ID NOs: 30 to 45.
[0309] A diagnostic method described in any of
[0301] to
[0308] , wherein the Inocle-related molecule is Inocle-α, and the expression level of the Inocle-related molecule is the sum of the expression levels of all nucleotides consisting of a base sequence represented by any of SEQ ID NOs: 30 to 45.
[0310] A diagnostic method described in any of
[0301] to
[0306] , wherein the Inocle-related molecule is at least one selected from the group consisting of Inocle-α, Inocle-β, Inocle-γ, and Inocle-δ, and the expression level of the Inocle-related molecule is the sum of the expression levels of all nucleotides consisting of a base sequence represented by any of SEQ ID NOs: 30 to 58.
[0152]
[0311] A diagnostic method described in any of
[0301] to
[0306] , wherein the Inocle-related molecule is at least one selected from the group consisting of Inocle-β, Inocle-γ, and Inocle-δ, and the expression level of the Inocle-related molecule is the expression level of a nucleotide consisting of at least one base sequence selected from the group consisting of base sequences represented by any of SEQ ID NOs: 46 to 58.
[0312] A diagnostic method described in any of
[0301] to
[0306] , wherein the Inocle-related molecule is at least one selected from the group consisting of Inocle-β, Inocle-γ, and Inocle-δ, and the expression level of the Inocle-related molecule is the sum of the expression levels of all nucleotides consisting of a base sequence represented by any of SEQ ID NOs: 46 to 58.
[0315] A diagnostic method described in any of
[0301] to
[0312] , wherein the expression level of the Inocle-related molecule is a value determined based on RPKM (reads per kilobase of contig per million total reads).
[0153]
[0316] A diagnostic method described in any of
[0301] to
[0312] and
[0315] , wherein the amount of Streptococcus Infantis is a value determined based on RPKM (reads per kilobase of contig per million total reads).
[0317] A diagnostic method described in any of
[0301] to
[0307] , wherein the Inocle-related molecule is Inocle-α, and the expression level of the Inocle-related molecule is the expression level of a protein consisting of at least one amino acid sequence selected from the group consisting of amino acid sequences represented by any of SEQ ID NOs: 1 to 16.
[0318] A diagnostic method described in any of
[0301] to
[0307] , wherein the Inocle-related molecule is Inocle-α, and the expression level of the Inocle-related molecule is the total expression level of all proteins consisting of an amino acid sequence represented by any of SEQ ID NOs: 1 to 16.
[0319] A diagnostic method described in any of
[0301] to
[0306] , wherein the Inocle-related molecule is at least one selected from the group consisting of Inocle-α, Inocle-β, Inocle-γ, and Inocle-δ, and the expression level of the Inocle-related molecule is the sum of the expression levels of all proteins consisting of at least one amino acid sequence selected from the group consisting of amino acid sequences represented by any of SEQ ID NOs: 1 to 29.
[0320] A diagnostic method described in any of
[0301] to
[0306] , wherein the Inocle-related molecule is at least one selected from the group consisting of Inocle-β, Inocle-γ, and Inocle-δ, and the expression level of the Inocle-related molecule is the expression level of a protein consisting of at least one amino acid sequence selected from the group consisting of amino acid sequences represented by any of SEQ ID NOs: 17 to 29.
[0154]
[0321] A diagnostic method described in any of
[0301] to
[0306] , wherein the Inocle-related molecule is at least one selected from the group consisting of Inocle-β, Inocle-γ, and Inocle-δ, and the expression level of the Inocle-related molecule is the sum of the expression levels of all proteins consisting of an amino acid sequence represented by any of SEQ ID NOs: 17 to 29.
[0155]
[0401] A cancer testing kit containing an inoculum-related molecule measurement reagent.
[0402] In addition, a cancer testing kit containing a Streptococcus Infantis detection reagent.
[0403] A cancer detection marker consisting of inocle-related molecules.
[0404] A marker consisting of inocle-related molecules used to determine whether or not a patient has cancer.
[0405] Inocle-related molecules used as diagnostic markers for cancer.
[0156]
[0501] A method for testing for immune diseases, comprising a step (C1) of measuring the amount of Streptococcus Infantis contained in the saliva of a subject.
[0502] Furthermore, a step (C2) of measuring the expression level of Inocle-related molecules contained in the saliva of the subject; and a step (C3) of calculating the expression level of the Inocle-related molecule per Streptcoccus Infantis by dividing the expression level of the Inocle-related molecule obtained in the step (C2) by the amount of Streptcoccus Infantis obtained in the step (C1); The inspection method described in
[0501] , which includes:
[0503] The immune disease is characterized by the expression of intermediate B cells, memory B cells, naive B cells, and CD14 + Monocytes and CD16 + An examination method described in claim
[0501] or
[0502] , wherein the immune disease is one in which monocytes are involved. [Example]
[0157] The present invention will now be described in more detail with reference to examples. It is not limited to:
[0158] (method) <Data Acquisition> The short-read sequence, long-read sequence, and closed Inocles contig are publicly available at DDBJ / GenBank / EMBL under accession number PRJDB18431. The sequence of InoC is available on Zenodo at DOI 10.5281 / zenodo.13138926.
[0159] <Collection of human saliva samples> This study was approved by the ethics committees of the University of Tokyo, the National Cancer Center Hospital East, and Sam Ratulangi University. Informed consent was obtained from all participants. For the development of the preNuc method, four saliva samples were collected from four Japanese participants (one from each group). For long-read metagenomics analysis, saliva samples were collected from 46 Japanese, 9 Indonesians, and 1 Thai. To estimate the detection rate of Inocles, saliva samples were collected from 68 Japanese and 20 Indonesians. In addition, a total of 45 saliva samples were collected from head and neck cancer (HNC) patients. Freshly collected saliva samples were immediately frozen at -80°C and stored until use.
[0160] <DNA extraction and purification from saliva samples> (Nuclease treatment) For the preNuc method, 1 mL of saliva sample was suspended in 500 μL of phosphate-buffered saline (PBS) and centrifuged at 12,000 × g for 10 minutes at 4°C. The supernatant was removed, and the resulting pellet was suspended in 500 μL of Tris / MgCl2 buffer (40 mM Tris / HCl, 10 mM MgCl2). DNase I (4 μL, 50 units) and RNase I (1 μL, final concentration 5 μg / mL) (Nippon Gene, Tokyo) were then added, and the mixture was incubated at 37°C for 30 minutes, followed by centrifugation at 12,000 × g for 10 minutes at 4°C. After removing the supernatant, DNA was extracted and purified as follows.
[0161] (DNA extraction) The pellet was mixed with ice-cold PBS (2 mL) and centrifuged (12,000 × g, 10 minutes, 4°C). The supernatant was discarded by decantation and the cells were collected. This procedure was repeated twice. The pellet was suspended in ice-cold TE20 (10 mM Tris, 20 mM EDTA, 800 μL). 50 μL of lysozyme solution (Sigma cat. L6876, 300 mg / mL) and 20 μL of achromopeptidase solution (FUJIFILM cat. 015-09951, 100 units / μL) were added and mixed on ice. The mixture was then incubated at 37°C for 2 hours with gentle shaking. 100 μL (final concentration 1%) of 10% SDS solution and 10 μL (1 mg) of Proteinase K solution (FUJIFILM cat. 161-28701, 0.1 mg / μL) were added, and the mixture was incubated at 55° C. for 1 hour with shaking.
[0162] (DNA purification) At room temperature, phenol / chloroform / isoamylalchol (equal volume, 1 mL) was added, and the tube was mixed by turning it upside down for 30 minutes, then centrifuged (12,000×g, 10 minutes, room temperature), and the supernatant was collected in a new 2 mL tube. Add 100 μL (final concentration = 0.3 M) of 3 M sodium acetate to the mixture on ice, mix, then add an equal volume of isopropanol (approximately 1 mL). Leave on ice for at least 10 minutes, then centrifuge (12,000 × g, 15–20 minutes, 4°C), and discard the supernatant by decantation. After decantation, the mixture was centrifuged again for approximately 1 minute, and the supernatant at the bottom was removed with a pipette. Add 1.5 mL of 75% EtOH to the pellet, vortex to thoroughly suspend the pellet, and centrifuge (12,000 × g, 5–10 minutes, 4°C). The supernatant was discarded by decantation. After decantation, the mixture was centrifuged again for approximately 1 minute, and the supernatant at the bottom was removed with a pipette. The pellet was dried under vacuum for 5 minutes, and then 100 μL of TE20 (10 mM Tris, 20 mM EDTA, 800 μL) was added to the pellet on ice to dissolve the DNA.
[0163] (RNase treatment-PEG precipitation) RNase (DNase-free) Solution (Nippongene cat. 313-01461) was added to the DNA solution to a final concentration of 10 μg / mL, mixed by inversion, and incubated at 37°C for 30 minutes. 60 μL of 20% PEG6000-2.5M NaCl (Hampton Research cat. HR2-533, HR2-637) was added at a volume 0.6 times the volume of the DNA solution, mixed well for several minutes, and then left on ice for 10 minutes. The mixture was centrifuged (12,000 × g, 15 minutes, 4°C), the supernatant was decanted, and then centrifuged again for approximately 1 minute. The supernatant was carefully removed with a pipette. The pellet was added with 75% EtOH (1.5 mL), vortexed, and centrifuged (12,000 × g, 5 minutes, 4°C). The supernatant was decanted and discarded. The pellet was centrifuged again for approximately 1 minute. The supernatant at the bottom was carefully removed with a pipette and rinsed with 75% EtOH. This 75% EtOH rinse was repeated twice. The pellet was dried under vacuum for 5 minutes, then TE (80–200 μL) was added and the pellet was left on ice for at least 30 minutes. It was then subjected to sequencing.
[0164] For DNA extraction from the small particle fraction in saliva, 1 mL of saliva sample was suspended in 1 mL of PBS and centrifuged at 5,000 × g for 10 min at 4°C. The resulting supernatant was filtered through a 0.45 μm pore size PVDF filter (Steriflip, Merck Millipore, Burlington, MA, USA). The filtrate was mixed with an equal volume of polyethylene glycol solution (20% PEG 6000 in 2.5 M NaCl) and stored at 4°C for 2 h. Small particles were then recovered by centrifugation at 20,000 × g for 45 min at 4°C. After removing the supernatant, the resulting pellet was suspended in 400 μL of TE20 buffer. Subsequent processing was performed as described above for DNA extraction, DNA purification, RNase treatment, and PEG precipitation.
[0165] Long-read metagenomic sequencing and data processing Library preparation was performed using the Ligation Sequencing Kit (SQK-LSK114) according to the manufacturer's protocol (Oxford Nanopore Technologies, Oxford, UK). Sequencing was performed using an Oxford Nanopore Technologies PromethION instrument and an R10.4.1 flow cell. Base calling was performed using dorado-0.6.2, and long reads with an average Q-score <20 were removed using NanoFilt (v2.8.0). After quality filtering, long reads were mapped to the T2T genome using minimap2 (2.26-r1175) with >90% sequence identity and >85% alignment coverage, and mapped reads were removed.
[0166] Short-read metagenomic sequencing and data processing Libraries were prepared using the TruSeq DNA Nano Kit (Illumina, San Diego, CA, USA) or TruSeq ChIP Kit (Illumina) according to the manufacturer's protocol. Sequencing was performed using an Illumina NovaSeq 6000. Quality filtering of NovaSeq reads was performed using fastp (v0.22.0) with the following options: n_base_limit 0 --trim_tail1 1 --cut_mean_quality 20 --qualified_quality_phred 20 --unqualified_percent_limit 50 --length_required 50 --detect_adapter_for_pe 2 --trim_poly_g --cut_tail --cut_tail_window_size 1 -dedup After quality filtering, the reads were mapped to the T2T genome using bowtie2 with default parameters, and the mapped reads were removed.
[0167] <Isolation of Streptococcus Strains Containing Inocle> To isolate Streptococcus strains with Inocle_004, a fresh saliva sample (sample ID: m_5) was plated on De Man-Rogosa-Sharpe (MRS) agar medium (Becton, Dickinson and Company USA) and cultured at 37°C for 1 day. The presence of Inocle_004 was confirmed by Sanger sequencing of the colony PCR products of five colonies using an Inocle_004-specific primer set (Inocle_004_forward: 5’-AACGCCAGCTCTTCTGGATA-3’ (SEQ ID NO: 200), Inocle_004_reverse: 5’-TGCTCGTGTAGGATCTGTCG-3’ (SEQ ID NO: 201)). Colony PCR was performed with 2×KAPA HiFi DNA polymerase (Kapa Biosystems, USA) and 10 μM primers for 35 cycles of 95°C for 3 minutes, 98°C for 20 seconds, 60°C for 30 seconds, and 72°C for 20 seconds. The bacteria of the five colonies obtained were cultured in MRS liquid medium at 37°C under aerobic conditions for 1 day. Genomic DNA was extracted from the culture using the DNeasy PowerSoil Pro Kit (QIAGEN, Germany). PCR for confirmation of the presence of Inocle_004 was performed using the DNA extracted from the liquid medium under the same conditions as colony PCR.
[0168] PCR for 16S rRNA was performed with V1-V2 region-specific primers (27Fmod and 338R) (10 μM each) and 2×KAPA HiFi DNA polymerase for 35 cycles of 95°C for 3 minutes, 98°C for 20 seconds, 60°C for 30 seconds, and 72°C for 20 seconds. An Agilent 2100 Bioanalyzer was used to confirm the size of the PCR products. Sequencing of the PCR products of the 16S rRNA gene was performed by the Sanger method. The obtained 16S rRNA sequences were aligned with the 16S rRNA sequences in the NCBI database to identify the most closely related species.
[0169] A whole-genome sequencing library of colony 5 was prepared using the NEBNext Ultra II DNA Library Prep Kit (New England Biolabs, USA). The genome sequence was obtained using an AVITI sequencer (Element Biosciences, San Diego, CA, USA). Sequences from AVITI were subsampled (0.5 million reads) and genome assembled using SPAdes with default parameters. The bacterial taxonomy of the isolates was determined using GTDB-tk (v2.3.2).
[0170] To calculate the ratio of Inocle-α to S. salivarius in native saliva, NovaSeq metagenomic reads from the m_5 saliva sample (from which Inocle_004 was isolated) were mapped to the isolated S. salivarius and Inocle_004 contigs using bowtie2 (v2.4.1). CoverM (v0.6.1) then calculated RPKM (reads per kilobase of contig per million total reads) from RPB (reads per base) obtained under the conditions --min-read-percent-identity 95 and --min-read-aligned-percent 85, and normalized by the total number of reads. Finally, the Inocle-α / S. salivarius ratio was calculated by dividing the Inocle-α RPKM value by the S. salivarius RPKM value.
[0171] <Long-read metagenomic assembly> Long read assembly was performed using Flye (2.9.2-b1786) with the --meta and --read-error 0.03 options. The circularity and sequencing depth of each contig were obtained from the metaFlye output file. Initially, two Inocles contigs (Inocle_008 and Inocle_018) were assembled as linear contigs. Reference-guided assembly was performed to verify the circularity of these two contigs. First, to refine the long reads corresponding to Inocle_008 and Inocle_018, which were assembled linearly, the long reads were mapped to the initially assembled contigs with a match rate of 95% or higher and alignment coverage of 20% or higher. The mapped long reads were then reassembled using metaFlye to obtain circular contigs.
[0172] To verify the accuracy of the long-read contigs, short reads were obtained from the two metagenomic DNA samples used to obtain the long reads. Short-read assembly was performed using MEGAHIT (v1.2.9). The short-read contigs were aligned to the long-read contigs using the -cx asm20 parameter in minimap2 (v2.26-r1175), and the similarity between the two was evaluated. All sequence statistics were obtained using Seqkit (v2.8.0).
[0173] <Classification of long-read contigs into microbial genetic elements> Small contigs less than 3 kb in length and contigs aligned to the T2T genome (>80% sequence identity and >20% alignment coverage) were removed from the long-read contigs with a sequencing depth of 10 or greater. To identify bacterial chromosome contigs, rRNA marker gene sequences in long-read contigs were predicted using Barrnap (v0.9) (https: / / github.com / tseemann / barrnap) and Prokka (v1.14.6). Other bacterial marker genes were predicted using fetchMG (v1.1), and contigs containing bacterial marker genes were classified as bacterial chromosome contigs. Binning analysis was performed using SemiBin2 (v2.0.2) with the parameters --environment human_oral and --sequencing-type=long_read. Contigs in bins containing bacterial chromosome contigs were classified as potential chromosome contigs. For non-bacterial chromosomal contigs, potential phage contigs were predicted using VirSorter2 (v2.2.4) with a score >0.9 and CheckV (v0.7.0) with a contamination <10%. For non-phage contigs, potential plasmid contigs were predicted using PlasClass (v0.1.1) with a score >0.7.
[0174] To identify contigs of eukaryotic virus and fungal origin, non-plasmid contigs were aligned to RefSeq virus genomes (downloaded July 6, 2023) and FungiDB (v64) using minimap2 (2.26-r1175) with the -cx asm20 parameter, >70% sequence identity, and >50% alignment coverage. Remaining contigs were left unclassified and aligned to the known plasmid database described in Reference B below and the integrated phage genome database described in References C, D, E, F, G, H, and I below.
[0175] Reference B: Galata, V., Fehlmann, T., Backes, C. & Keller, A. PLSDB: A resource of complete bacterial plasmids. Nucleic Acids Res. 47, D195-D202 (2019). Reference C: Nayfach, S. et al. Metagenomic compendium of 189,680 DNA viruses from the human gut microbiome. Nature Microbiology 6, 960-970 (2021). Reference D: Gregory, A. C. et al. The Gut Virome Database Reveals Age-Dependent Patterns of Virome Diversity in the Human Gut. Cell Host Microbe 28, 724-740.e8 (2020). Reference E: Camarillo-Guerrero, L. F., Almeida, A., Rangel-Pineros, G., Finn, R. D. & Lawley, T. D. Massive expansion of human gut bacteriophage diversity. Cell 184, 1098-1109.e9 (2021). Reference F: Nishijima, S. et al. Extensive gut virome variation and its associations with host and environmental factors in a population-level cohort. Nat. Commun. 13, 1-14 (2022). Reference G Yuya, K., Nishijima, S., Kumar, N., Hattori, M. & Suda, W. Long-read metagenomics of multiple displacement amplified DNA of low-biomass human gut phageomes by SACRA pre-processing chimeric reads. DNA Research 28, dsab019 (2021) Reference H Li, S. et al. A catalog of 48,425 nonredundant viruses from oral metagenomes expands the horizon of the human oral virome. iScience 25, 104418 (2022). Reference I Roux, S. et al. IMG / VR v3: An integrated ecological and evolutionary framework for interrogating genomes of uncultivated viruses. Nucleic Acids Res. 49, D764-D775 (2021)
[0176] ORFs in contigs were predicted with the default parameters of Prokka (v1.14.6). All ORFs were aligned to the UniRef90 database using DIAMOND (v2.1.8) with the conditions of --sensitive and e-value < 1e-10. Contigs containing more than 50% ORFs that did not match the UniRef90 database were classified as strong candidates for unrecognized genetic factors. For clustering of candidate contigs, clustering based on average amino acid sequence identity (AAI) was performed. For all ORFs of 85 strong candidate contigs of unrecognized genetic factors, "all vs. all aligned" was performed using DIAMOND with the conditions of --evalue 1e-5, --max-target-seqs 10000, --query-cover 50, and --subject-cover 50. Clustering analysis was performed for contig pairs that met the criteria of more than 80% shared ORFs, more than 30% average AAI, and more than 10 shared ORFs. Clustering was performed using MCL (v14-137) with the parameters of -te 8, -I 2.0, and --abc. Contigs that were more than 1.5 times smaller or larger than the average contig size within the same cluster were treated as separate clusters. Long-read mapping was performed using the map-ont option of minimap2 (2.26-r1175). The mapping results were visualized by Integrative Genomics Viewer (IGV).
[0177] For calculating RPKM, first, CoverM (v0.6.1) (https: / / github.com / wwood / CoverM) was used to obtain the reads per base (RPB) with the conditions of --min-read-percent-identity 95 and --min-read-aligned-percent 85. Next, RPB was normalized by the total number of reads in each sample to calculate RPKM.
[0178] <Genome analysis of Inocles> Based on the cumulative GC Skew, the replication origin and terminus were predicted using iRep (v1.10). Genome visualization was performed using Proksee. Insertion sequences were predicted using ISEscan (v1.7.2.3) with default parameters. Among all high-quality circular contigs, those having an ORF aligned to at least one InoC gene with an e-value < 1e-10 using DIAMOND were identified as additional Inocles contigs. The average nucleotide identity (ANI) was calculated using Pyani (v0.2.12) with the -m ANIb option. The PF of Inocles was obtained using Roary (v3.13.0) with the -i 50 option to perform ORF clustering. Based on the similarity of the PF repertoires, a dendrogram was created using the ward.D2 method with the Euclidean distance calculated from the presence / absence information of PF. Short-read contigs assembled from Inocles contig reconstruction samples using MEGAHIT were aligned to Inocles contigs using minimap2 (2.26-r1175), and the fragmentation positions in the short-read contigs were determined with the conditions of >99% sequence identity and an alignment length of 500 bp or more.
[0179] <Prediction of the host bacteria of Inocles> To identify the ORFs of Inocles that are homologous to those of the host bacteria, all Inocles ORFs were aligned against the UniRef90 database using DIAMOND (v2.1.8), and the results of the best hits were extracted under the condition of e-value < 1e-10 and used for subsequent analysis. The number of Inocles genes showing homology to those of the host bacteria was calculated by summing the number of genus-level classifications on UniRef90 corresponding to the aligned Inocles genes.
[0180] The tetranucleotide frequency was obtained from the chromosomal, plasmid, and phage-derived sequences of Inocles and Streptococcus using CheckM (v1.1.3). The relative synonymous codon usage was obtained using Codon Usage Generator (ver. 2.4) (http: / / bioinfo.ie.niigata-u.ac.jp / ?Codon+Usage+Generator). All chromosomal, plasmid, and phage sequences of the Streptococcus genus were obtained from the RefSeq database. Dendrograms based on the similarity of tetranucleotide frequency and codon usage were created using the ward.D2 method based on Manhattan distance.
[0181] Short reads of the bacterial pellet fraction and 0.45 μm filtered fraction obtained from saliva samples were mapped against the Inocles contigs obtained from the same samples. Mapping was performed using bowtie2 (v2.4.1), and the reads per kilobase per million reads (RPKM) was calculated by normalizing the reads per base (RPB) obtained under the conditions of --min-read-percent-identity 95 and --min-read-aligned-percent 85 using CoverM (v0.6.1) with the total number of reads in each sample.
[0182] <Functional gene annotation of Inocles> InoC was identified by ORF clustering (-c 0.7, n 5, aS 0.5) using cd-hit (v4.8.1). The three-dimensional structure of InoC was predicted using ColabFold (v1.5.5). The phylogenetic tree of InoC was constructed using FastTree (v2.1.11) after multiple alignment using MAFFT (v7.505) and visualized using iToL80. ORF functional annotation was performed using Prokka (v1.14.6). Mobilization-related genes were predicted using ICEscreen (v1.2.0) with default parameters. ORF subcellular localization prediction was performed using Psort (v3.0) with the --positive option. KEGG orthology (K numbering) was assigned to ORFs using eggNOG mapper (v2.1.11) (--sensmode ultra-sensitive), and a KEGG BRITE hierarchy was constructed based on the resulting K numbers. Pfam domain and active site predictions for signal peptidases were performed using InterProScan. Transmembrane regions in cell wall / membrane-associated ORFs were predicted using TMHMM-2.0 with default parameters, and the 2D structures of proteins on membranes were visualized using protter. LPXTG-like motif predictions were performed using CW-PRED with default parameters.
[0183] Homologous genes between Inocles and Streptococcus chromosomes were identified by performing a similarity search between Inocles ORFs and Streptococcus chromosome ORFs using DIAMOND (v2.1.8) with the options --sensitive and e-value < 1e-10.
[0184] <Sequence analysis of submitted datasets> The deposited metagenomic sequences in HMP were downloaded from the NCBI SRA server. Low-quality reads were removed using fastp with the following options: --n_base_limit 0 --trim_tail1 1 --cut_mean_quality 20 --qualified_quality_phred 20 --unqualified_percent_limit 50 --length_required 50 --detect_adapter_for_pe 2 --trim_poly_g --cut_tail --cut_tail_window_size 1--dedup
[0185] After quality filtering, reads were mapped to the T2T genome using bowtie2 (default parameters), and unmapped reads were extracted as filter-passed reads. Samples with filter-passed reads exceeding 4 million reads were used for subsequent analysis. Filter-passed short reads were mapped to non-redundant InoC genes clustered by 100% sequence identity using bowtie2 (v2.4.1). RPKM was calculated by normalizing the reads per base (RPB) obtained using CoverM (v0.6.1) with --min-read-percent-identity 95 and --min-read-aligned-percent 85 by the total number of reads for each sample. RPKM for each Inocles taxon was calculated by summing the RPKM for all InoC genes belonging to that taxon.
[0186] To confirm transcription of Inocles ORFs, metatranscriptome short reads were mapped to the base sequences of Inocles PFs using bowtie2 (v2.4.1), and RPKM was calculated using CoverM (v0.6.1) by normalizing the RPB obtained under the conditions of --min-read-percent-identity 95 and --min-read-aligned-percent 85 by the total number of reads in each sample.
[0187] Genomic sequences (Illumina reads) from oral bacteria were obtained by searching the NCBI database using the keywords "oral," "isolate," and "bacteria" and excluding metagenomic sequences. These short reads were mapped to the InoC gene using bowtie2 (v2.4.1).
[0188] To compare the abundance of inocles, saliva metagenomic sequences from healthy individuals (HC) and cohorts with pancreatic ductal adenocarcinoma (PDAC), rheumatoid arthritis (RA), and colorectal cancer (CRC) were collected from the NCBI SRA server. One sample was randomly selected from longitudinally collected participants in the CRC study. Quality filtering and removal of human reads were performed using the same procedures as in the HMP study. Samples with filter-passed reads exceeding 4 million reads were used for comparative analysis.
[0189] <Single-cell RNA-seq analysis> The dataset of single-cell populations was selected based on the work of Kashima et al. (Ref. J), which describes the details of the experimental method and bioinformatics processing. Yukie K., et al. Construction and Characterization of a Single-Cell Catalog of Aging Immune Cells. Under submission. (2024).
[0190] Briefly, peripheral blood mononuclear cells (PBMCs) were subjected to library preparation using 10x Genomics Chromium Next GEM Single Cell Multiome ATAC + Gene Expression according to the manufacturer's user guide (CG000365 Rev A, CG000337 Rev A, 10x Genomics). scRNA-seq was performed on a NovaSeq 6000 (Illumina). The resulting sequencing data was processed using Cell Ranger ARC (Cellranger-arc-2.0.0) (default settings, intron-inclusive mode, reference: GRCh38-2020-A-2.0.0). Subsequently, NovaSeq reads were mapped to hg38 using STAR, and human gene count data were obtained using Seurat (v4.1.1). Count data were normalized and integrated using Harmony (v0.1.0) to minimize batch effects. Cell clustering was performed using Seurat, and clusters were manually annotated based on canonical markers listed in the Azimuth reference. GO (Gene Ontology) analysis was performed using ShinyGO 8.0.
[0191] <Proteome analysis of human plasma samples> Blood samples and plasma fractions were collected for proteomic analysis using Olink. Antibody reactions targeting 5,420 proteins were performed in all samples according to the manufacturer's protocol (Olink Proteomics, Uppsala, Sweden) without bridging between different experimental reactions. Sequencing libraries were prepared according to the manufacturer's protocol (Olink Explore HT), followed by sequencing on a NovaSeq 6000 (Illumina). The resulting sequence counts were converted to normalized protein expression values (NPX) and subjected to analysis.
[0192] <Associative analysis among Inocle-α, S. salivarius, blood cells, and plasma proteome> The abundance of Inocle-α was obtained as RPKM, and the abundance of the genus Streptococcus was obtained using MetaPhlAn4 (v4.1.0). The correlation analysis was performed based on the partial Spearman correlation coefficient using the psych package on RStudio (v2023.12.0+369), considering confounding factors such as age and gender.
[0193] <Ratio of Inocle-α to S. infantis> To calculate the ratio of Inocle-α to S. infantis, the HQ-MAG of S. infantis was identified from the original reconstructed HQ-MAG (high quality metagenome assembled-genomes) by long-read assembly and binning using GTDB-tk (v2.3.2). The HQ-MAG of S. infantis was dereplicated using dRep (v3.4.3) with the options --S_algorithm fastANI --P_ani0.9 --S_ani0.95 --coverage_method bigger --cov_thresh0.85.
[0194] Short reads were mapped to the S. infantis HQ-MAG using bowtie2 (v2.4.1), and RPKMs were obtained using CoverM (v0.6.1) with the --min-read-percent-identity 95 and --min-read-aligned-percent 85 options. The total amount of S. infantis was obtained by summing all RPKMs. To obtain the composition of inocles, short reads were mapped to non-overlapping InoC genes using bowtie2 (v2.4.1), and RPKMs were obtained using CoverM (v0.6.1) with the --min-read-percent-identity 95 and --min-read-aligned-percent 85 options. The RPKM of each Inocle-α was obtained by summing all RPKMs of the InoC genes of each Inocle-α. The RPKM ratio between Inocle-α and S. infantis was then calculated.
[0195] <Statistical analysis> All statistical analyses were performed using RStudio (v2023.12.0+369). A two-tailed Student's t-test was used to compare preNuc-treated and untreated groups. Fisher's exact test was used to compare the detection rates of Inocle-α and Streptococcus between HNC and HC groups. For the association analysis between Inocle-α and human diseases, MaAsLin2 (v1.0.0) was used. Statistical tests were performed with default parameters, excluding specific random effects. Data for HC and HNC patients were analyzed using MaAsLin2, with sex, age, smoking status, and cancer treatment status specified as random effects. Similarly, for the comparison of PDAC, RA, and CRC patients with HC, sex and age were used as random effects. Statistical results were considered significant when p values were <0.05.
[0196] Example 1: Preparation of a nucleic acid solution derived from oral bacteria contained in a saliva sample Long-read metagenomic sequencing optimized for human saliva samples Long-read sequencing improves metagenomic assembly for reconstructing ECE with high completeness. However, because 85-95% of the DNA in human saliva is human genomic DNA, it is difficult to obtain sufficient sequencing depth from microorganisms contained in human saliva. Therefore, we developed a simple method to extract high-molecular-weight DNA with reduced human genomic DNA from saliva samples.
[0197] After collecting the bacterial pellet by centrifugation, the bacterial pellet was treated with nuclease under conditions optimal for nuclease activity to remove extracellular DNA, which was thought to be primarily human DNA. This treatment method is called "preNuc." After preNuc treatment, bacteria were enzymatically lysed and high-molecular-weight DNA was extracted (Figure 1A). Notably, preNuc treatment reduced the average percentage of human DNA in short-read sequences prepared from fresh and frozen saliva samples from 90% to 40% (Figure 1B and Table 9).
[0198] Using the Oxford Nanopore Technologies PromethION sequencer, preNuc treatment resulted in a two-fold increase in the total amount of microbial bases compared to preNuc-untreated samples (Figure 1C). Furthermore, there was no reduction in read length (Figure 1D). Furthermore, the Spearman's rank correlation coefficient (ρ) for bacterial composition at the species level averaged 0.83 ± 0.08 between preNuc-treated and untreated saliva samples (Figure 1E). These results demonstrate that preNuc can easily reduce human DNA in saliva samples and extract high molecular weight DNA with minimal impact on bacterial composition.
[0199] We then obtained PromethION sequences from 46 Japanese saliva samples, yielding an average of 3.2 million high-quality, non-human read sequences (23.4 Gb per sample, N50 length 12.5 kb) (Table 10). De novo metagenomic assembly using metaFlye (Table 11) was performed to obtain contigs. From these contigs, we excluded human genome-derived contigs, low-quality contigs (depth less than 10), and small contigs (less than 3 kb). These contigs were used for further analysis as high-quality contigs, because they showed an average similarity of 98.95% with the corresponding short-read contigs assembled using NovaSeq sequences (Figure 7).
[0200] [Example 2] Identification of a novel ECE To deepen our understanding of the genetic diversity of ECEs in the human oral microbiome, we attempted to identify novel ECEs, as shown in Figure 8A. First, bacterial chromosome contigs were identified from high-quality contigs based on the presence of bacterial rRNA marker genes and fetchMG-defined marker genes. Contigs belonging to the same bin as a bacterial chromosome contig were identified as potential bacterial chromosome contigs. Non-bacterial chromosome contigs with a virsorter2 score >0.9 and contamination (CheckV) <10% were identified as phage contigs. Among non-phage contigs, those with a PlasmClass score >0.7 were identified as plasmid contigs. Eukaryotic viral and fungal contigs were identified based on alignments with reference genomes. Finally, contigs that were not identified as bacterial, phage, plasmid, viral, or fungal contigs were designated as unclassified contigs.
[0201] As a result, 69,881 high-quality contigs that were not aligned to the human genome were classified into known genetic elements: bacterial contigs, phage contigs, plasmid contigs, viral contigs, and fungal contigs (Fig. 8A), of which 85.8% were derived from bacterial chromosomes, 1.5% from phages, 2.9% from plasmids, and 9.8% were unclassified contigs (Fig. 2A).
[0202] Of the 6,852 unclassified contigs, 85 were considered to be unrecognized genetic elements, with many genes whose functions were unknown (Fig. 2A). Specifically, more than 50% of the genes contained in these 85 contigs did not match any of the UniRef90 database, a figure lower than that for known microbial genetic elements (Fig. 8B). For these 85 contigs, we clustered contigs that shared 80% or more of their genes and had an average amino acid identity of over 30% within the cluster, resulting in 19 contig clusters (clusters 0001–0019) that contained no singletons (Figure 2B). We quantified the contig content of each cluster by long-read mapping, and cluster 0001, consisting of 13 contigs, showed the highest contig content and detection rate (Figure 2B).
[0203] Within cluster 0001, 8 of the 13 contigs were ranked in the top 50% of all contigs. Furthermore, one of the contigs in cluster 0001, sample ID: OO1116B8239, was among the top 20 most abundant contigs. This suggests that cluster 0001 is a genetic element abundant in human saliva (Figure 2C). None of the contigs in cluster 0001 could be aligned to plasmid or phage databases with greater than 95% identity and greater than 85% coverage. These results indicate that cluster 0001 constitutes a potentially novel genetic element that is widespread in the human oral microbiome.
[0204] The average contig size of cluster 0001 was 352 kb (range, 293–395 kb) (Figure 2D), with an average number of open reading frames (ORFs) of 313 (Figure 2E). The morphology of 13 contigs was identified as circular based on the metaFlye assembly. Long-read mapping revealed uniform sequence depth throughout the contigs, with no missing regions. This suggests that these contigs are unlikely to have been artificially misassembled (Figure 9). Notably, there was no accumulation of alignment start and end positions within the contigs. This is characteristic of linear contigs with terminal direct repeats, confirming the circular morphology of these contigs (Figure 9). Furthermore, GC skew analysis identified potential replication origins and termini (Figure 2F), revealing that the orientation of ORFs at the replication origin and termini was reversed (Figure 2G and Figure 10). This suggests that this cluster is a circular replicon.
[0205] We discuss why these genetic factors have been overlooked in previous metagenomic studies. One possible reason for this is the low sequencing depth in previous studies, but in our study, preNuc increased the long-read sequence coverage of cluster 0001 within a specific sample from 9.6x to 30x, contributing to the reconstruction of high-quality contigs (Figure 2H). Another reason is the difficulty of assembling multiple regions using short reads. Although some contigs were successfully reconstructed by short read assembly, these assemblies were fragmented, with 48% (48 out of 102) of the breakpoints overlapping with repeated insertion sequences (IS) (Figures 2I and 11). It was also confirmed that 17% (17 out of 102) of the breakpoints overlapped with other hypervariable regions (Figures 2I and 11). However, in the present invention, this difficulty was overcome by using long reads.
[0206] We named this novel genetic element "Inocle" (insertion sequence encoding; oral origin; circle genomic structure). Considering the lack of bacterial marker genes in Inocles, Inocles is likely to be a novel ECE in the human oral microbiome.
[0207] [Example 3] Search and identification of Inocles marker genes Next, we investigated the extent of undiscovered sequences within the Inocles family. First, we searched for marker genes that could serve as an efficient strategy for distinguishing Inocles contigs from other metagenomic contigs. To identify such marker genes, we clustered all Inocles ORFs and identified a single gene that met the following criteria: a gene with an amino acid sequence identity of 70% or more and alignment coverage of 50% or more across all initially identified Inocles contigs. We named this gene "InoC" (Inocle conserved gene).
[0208] InoC is a single-copy gene with an average length of 454 amino acids and contains a replication relaxation domain (Pfam ID: PF13814) in the N-terminal domain (Figure 3A). No genes similar to InoC were found in the NCBI nucleotide database, suggesting that InoC is an Inocles-specific gene.
[0209] Using InoC, we performed an extended analysis to identify additional Inocles family sequences from high-quality circular contigs obtained from saliva samples of 46 Japanese subjects (Figure 3A). Furthermore, to further investigate the global distribution of Inocles family contigs, we applied the same strategy to long-read contigs assembled from saliva samples of nine Indonesian and one Thai healthy subjects (Tables 10 and 11).
[0210] As a result, 16 additional circular contigs encoding InoC were identified (Fig. 3A), all of which contained candidate replication origins and termination sites (Fig. S12). Finally, 29 closed circular Inocles contigs were identified, including 21 contigs from 17 Japanese subjects, 7 contigs from 4 Indonesian subjects, and 1 contig from 1 Thai subject (Table S12).
[0211] Example 4: Four Inocles taxa and their genetic and ecological differences A phylogenetic tree of the InoC genes of 29 Inocles contigs revealed four distinct groups, designated Inocle-α, β, γ, and δ (Fig. 3B). The average amino acid identity of InoC within a group was 94.5% and between groups was 64.1% (Fig. S13A). Inocle-α (average 368 kb) was the largest contig, while Inocle-δ (average 245 kb) was the smallest (Fig. S13B).
[0212] The 29 Inocles contigs were designated Inocle001 to 029. Inocles are classified into Inocle-α consisting of 16 types, Inocle001 to 016 (SEQ ID NOs: 149 to 164, respectively), Inocle-β consisting of 6 types, Inocle017 to 022 (SEQ ID NOs: 165 to 170, respectively), Inocle-γ consisting of 4 types, Inocle023 to 026 (SEQ ID NOs: 171 to 174, respectively), and Inocle-δ consisting of 3 types, Inocle027 to 029 (SEQ ID NOs: 175 to 177, respectively).
[0213] The average nucleotide identity within each group was 88.9% (Figure 13C), and the alignment coverage between groups was insufficient (>20%). The number of tRNAs varied among groups (Figure 13D). The GC content of Inocle-α, β, and γ was 32% on average, while that of Inocle-δ was 24% on average (Figure 13E). All Inocles contained ISs, but Inocle-α was particularly IS-rich (average of 9 ISs per contig) (Figure 13F). This suggests that Inocles expand their genetic diversity through mobile genetic elements called ISs.
[0214] A total of 1,619 protein families (PFs) with amino acid sequence identity >50% were identified across the Inocles contigs. Of these, 1,537 PFs (95%) were found exclusively in specific groups (Figure 3C). Genetic differences among these four groups suggest that these four Inocles taxa have distinct evolutionary histories.
[0215] To investigate the geographic distribution of each Inocles taxon, we downloaded saliva shotgun metagenomic sequences from healthy individuals in China (n = 47), the Philippines (n = 24), France (n = 16), the United States (n = 64), and Fiji (n = 237) from public databases (Table 13). We also obtained short-read metagenomic sequences from saliva samples of 68 healthy Japanese and 20 healthy Indonesian individuals (Table 13). We identified Inocles-positive samples by mapping short reads to the InoC gene, and found that an average of 74% of individuals were Inocles-positive (Figure 3D). Indonesia had a higher rate of Inocles detection (90%) than other countries, while Japan had the lowest Inocles detection rate (64%) (Figure 3D). In all countries, Inocle-α was the predominant group among each Inocles taxon, with an average detection rate of 57%, but non-industrialized countries such as Indonesia, the Philippines, and Fiji had a higher detection rate of Inocle-β than the other countries (Figure 3D).
[0216] Next, we mapped the InoC gene to shotgun metagenomic sequences from various sites collected by the Human Microbiome Project (HMP) to clarify the localization of each Inocles taxon in the human body. Each Inocles taxon was detected in the oral cavity, but only in small amounts in feces, the external nares, the posterior vaginal vault, and the postauricular groove (Figure 3E). Within the oral cavity, Inocle-α was primarily detected on the dorsum of the tongue, Inocle-β and Inocle-δ on the buccal mucosa, and Inocle-γ on the gingiva (Figure 3E). These results suggest that the predominant Inocles taxon differs depending on the oral site.
[0217] Example 5: Inocles are large plasmid-like elements in Streptococcus. Bacterial cells and small extracellular particles such as phages can be separated by centrifugation followed by 0.45 μm filtration into a centrifugal pellet fraction (bacterial cells) and a 0.45 μm filtration fraction (extracellular particles). Fractionation of saliva samples revealed that the number of Inocles-derived sequences in the centrifugal pellet fraction was on average 16-fold higher than that in the 0.45 μm filtration fraction, suggesting that Inocles are intracellular ECEs (Figure 4A).
[0218] Some intracellular ECEs have homologous genes to their host bacteria. Therefore, to identify the host bacteria of Inocles, we performed a homology search using Inocles genes against all bacterial genera registered in the UniRef database. As a result, an average of 53 ORFs per contig showed homology to Streptococcus genes (Fig. 4B), suggesting that Streptococcus is a potential host bacterium for Inocles.
[0219] Because Inocles are thought to be intracellular ECEs of oral bacteria, we investigated the genome dataset of oral bacterial isolates, including sequences derived from Inocles, deposited in the NCBI database. However, after searching over 3,000 oral bacterial isolates, we found no sequences that mapped to the InoC gene.
[0220] Therefore, we attempted to isolate and culture a bacterial strain containing only one inoculus (Inocle_004). First, salivary bacteria were cultured on De Man-Rogosa-Sharpe (MRS) agar medium (Figure 4C). Colony PCR was performed on five colonies using Inocle_004-specific primers, and three colonies whose PCR product sequences perfectly matched the Inocle_004 contig were identified (Figure 4C and Table 14).
[0221] Next, to obtain sufficient genomic DNA for whole-genome sequencing, the strain (colony 1) was grown in MRS liquid medium and high-depth short-read sequencing was performed from the culture medium. Although a sequence depth of over 2,500 was obtained, none of the sequences derived from Inocle_004 were included, suggesting that Inocle_004 was lost from the host cells during liquid culture. To confirm the loss of Inocle_004 during liquid culture in other isolates, PCR was performed using Inocle_004-specific primers on liquid-cultured colonies 2 and 3, but no PCR product was obtained (Figure 4C). These results confirm that Inocle_004 was lost from host cells during cultivation in liquid medium, and this feature may be the reason why Inocles was not detected in the published genome sequence.
[0222] Whole genome sequences (colony 1 strain) and 16S rRNA gene sequences (colonies 2 and 3) of isolates grown in liquid culture revealed that all strains were Streptococcus salivarius (Fig. 4C, Table 14), suggesting that S. salivarius is one of the candidate host bacterial species for Inocle_004.
[0223] To estimate the copy number of Inocle_004 in the potential host bacteria in saliva, we mapped the metagenomic reads from saliva to the S. salivarius genome and Inocle_004 contig of colony 1, and the estimated ratio of Inocle_004 to S. salivarius was 1:15.
[0224] Next, we compared the genome signatures of Inocles with those of Streptococcus chromosomes, plasmids, and phages. The tetranucleotide frequencies of Inocles were relatively similar to those of plasmids, rather than to those of chromosomes or phages. However, the dendrogram obtained from the tetranucleotide frequency similarity showed that Inocles formed a unique cluster (Fig. 4D). The same trend was observed for codon usage (Fig. 4E), revealing that Inocles exhibit a distinct nucleotide composition and codon usage frequency from known plasmids.
[0225] We identified VirB4 and VirD4 in all Inocles contigs, which are required for the transfer of plasmids from donor cells to recipient cells via the type IV secretion system. Thus, Inocles is considered a novel large plasmid-like element in Streptococcus, potentially capable of horizontal transfer.
[0226] [Example 6 Gene annotation of Inocles] To clarify the biological roles of Inocles in the host bacterium, we performed functional annotation of 1,619 protein families (PFs) encoded by the Inocles contig. First, we examined the subcellular localization of these PFs and identified 644 (40%) PFs as cytoplasmic proteins and 620 (38%) as proteins with unknown subcellular localization (Figure 5A). Among the 644 cytoplasmic proteins, 21 functional genes were annotated using Prokka.
[0227] Of the 11 cytoplasmic genes conserved across all Inocles taxa, six were annotated as genes related to DNA repair and recombination (ko03400) based on the Kyoto Encyclopedia of Genes and Genome (KEGG) BRITE database, all of which have DNA damage repair-related functions, such as exonuclease, DNA gyrase, and DNA polymerase III (Figure 5B).
[0228] Among genes whose subcellular localization was unknown, we identified a RecD-like DNA helicase, which was also associated with genes involved in DNA damage repair (Fig. 5C). We also identified other stress response-related genes detected in multiple Inocles taxa, such as RpoS, which contributes to stress responses to multiple stressors, and NrdH, which confers oxidative stress resistance. Based on metatranscriptome sequencing of human saliva, we confirmed the transcriptional activity of six of the seven DNA damage repair genes, RpoS, and NrdH (Figure 5B and C).
[0229] Next, we focused on 329 PFs (20%) that were potentially cell wall / membrane-associated proteins (Figure 5A). Among them, we identified eight functionally annotated genes, two of which were conserved across all Inocles taxa, one of which was classified as a peptidoglycan biosynthesis / degradation protein based on the KEGG BRITE database (Figure 5D). This protein possesses a transmembrane domain at its N-terminus and a transglycosylase domain and a penicillin-binding protein transpeptidase domain in its extramembrane region (Figure 5E). Therefore, this protein is a potential bifunctional transglycosylase that catalyzes the polymerization and cross-linking of glycan chains in peptidoglycan biosynthesis.
[0230] Another cell wall / membrane-related gene was annotated as sortase A (Figure 5D). This protein has a transmembrane domain at the N-terminus, and the sortase domain is located in the extracellular region (Figure 5F). This protein is thought to recognize the LPXTG-like motif in cell wall proteins and anchor them to peptidoglycan.
[0231] Six LPXTG-like motif PFs predicted to be recognized by sortase A were also identified. All Inocles taxa encoded at least one LPXTG-like motif PF (Figure 14A). These LPXTG-like motif PFs have two transmembrane domains spanning the extracellular region, one at each of the N-terminus and C-terminus (Figure 5G and Figure 14B). It was revealed that the LPXTG-like motif has a signal peptide for membrane translocation at the C-terminus of the extracellular region and a signal peptidase cleavage site at its N-terminus (Figure 5G and Figure 14B). Furthermore, signal peptidases from Inocle-α, β, γ with a signal peptidase S26 domain (PF10502) (Figure 14C) and a lysine active site for cleaving the signal peptide were identified (Figure 5H and Figure 14D). As a result of confirming the transcriptional activity of these cell wall-related genes, their bioactivity in the natural oral environment was suggested (Figure 5D, Figure 14A, and Figure 14C).
[0232] Next, it was investigated whether Inocles functionally overlaps with the Streptococcus chromosome. This is a phenomenon commonly observed between plasmids and the host bacterial chromosome. On average, 20 ORFs (6%) per Inocle were identified as homologous genes between the Streptococcus chromosome and Inocles (Figure 15). These genes included RecD-like DNA helicase, DNA polymerase III subunit α, bifunctional transglycosylase, and sortase A (Figure 5I).
[0233] <Inocle-α is correlated with the overall immune status of humans> Next, we investigated the relationship between Inocles and human physiological functions. First, we investigated whether Inocle-α was associated with the gender and age of the host (healthy individuals), but no significant association was found (Table 15, Figures 16A and 16B). Because the detection rates of Inocle-β, γ, and δ in the Japanese sample population were low (Figure 3D), we were unable to statistically analyze the data for Inocle-β, γ, and δ.
[0234] Furthermore, we investigated the association using a human multi-omics dataset. Peripheral blood mononuclear cells (PBMCs) were collected from 29 subjects who provided saliva samples in this study, and single-cell RNA sequencing (scRNA-seq) analysis was performed (Table 16). PBMC cell populations were classified into 19 cell types based on the Azimuth reference. Spearman's partial correlation coefficient, taking into account age and sex as confounding factors, showed a significant positive correlation between Inocle-α and intermediate B cells, memory B cells, and naive B cells, but not plasma cells (Figure 6A). + Monocytes and CD14 + A significant negative correlation was observed between Streptococcus and CD16. + Monocytes and CD14 + No interaction was observed between monocytes (Fig. 6B).
[0235] Because there was a positive correlation between some B cells and Inocle-α, we further analyzed gene expression in B cells. We observed that the expression of 687 genes was significantly increased and the expression of 62 genes was significantly decreased in the Inocle-α-positive group compared to the Inocle-α-negative group. Gene Ontology (GO) analysis revealed that 44 pathways were enriched in the Inocle-α positive group compared to the Inocle-α negative group (>1.4-log10 FDR [<0.05 FDR]) (Figure 6B). Among these 44 pathways, cellular senescence was the most enriched pathway in the Inocle-α positive group (Figure 6B). Furthermore, five pathways involved in the response to bacterial and viral infections, namely hepatitis B, Escherichia infection, human T-cell leukemia virus type 1 infection, Kaposi's sarcoma-associated herpesvirus infection, and human cytomegalovirus infection, were also enriched (Figure 6B).
[0236] To further examine the association between Inocle-α and the human systemic immune system, expression data of a total of 5,420 plasma proteins were analyzed using Proximity Extension Assay technology (Olink) from 40 healthy Japanese subjects. In partial Spearman correlation considering age and gender as confounding factors, 1,187 and 62 proteins were significantly positively and negatively correlated with Inocle-α, respectively. GO enrichment analysis revealed that 15 pathways were particularly enriched in the Inocle-α positive group (>5-log10 FDR), four of which were related to the response to bacterial and viral infections such as epithelial cell signaling in pathogenic Escherichia coli infection, Escherichia infection, hepatitis B, Salmonella infection, Kaposi's sarcoma-associated herpesvirus infection, and Helicobacter pylori infection (Figure 6B), and this result was partially consistent with the transcriptional differences observed in the B cell population. Furthermore, it was revealed that in participants rich in Inocle-α, the expression levels of proteins related to lipid and atherosclerotic arteriosclerosis, endocytosis, TNF signaling, NF-κB signaling pathway, and induction of apoptosis were high (Figure 6B).
[0237] <Correlation between Inocle-α and cancer - 1> Furthermore, Inocle-α was found to be associated with several cancer-related pathways in the plasma proteome, including proteoglycans in cancer, chronic myeloid leukemia, choline metabolism in cancer, pancreatic cancer, and bladder cancer (Figure 6B). Therefore, we investigated the possible association between Inocle-α and cancer. Saliva samples were collected from 45 head and neck cancer (HNC) patients (22 of whom were treatment-naive), and metagenomic short reads were obtained (Table 15). The abundance of Inocle-α was compared between 53 healthy controls (HC) and 45 HNC patients, taking into account age, sex, smoking status, and cancer treatment information as confounding factors (Table 15).
[0238] The detection rate of Inocle-α was significantly lower in the HNC group (27%) than in the HC group (60%), and the abundance of Inocle-α in the HNC group was also lower than in the HC group (Fig. 6C). However, the detection rate and abundance of Streptococcus were not significantly reduced in the HNC group compared with the HC group (Fig. 6D). These results suggest that host bacteria carrying Inocle-α are selectively reduced in the saliva of HNC patients.
[0239] We further analyzed whether the association of Inocles was observed in other cancer types. To this end, we downloaded publicly available salivary metagenomic short reads from patients with pancreatic ductal adenocarcinoma (PDAC) and colorectal cancer (CRC), as well as healthy controls (HCs) (Table 17). As a reference for inflammatory diseases, we also obtained data from patients with rheumatoid arthritis (RA) (Table 17). Statistical analysis revealed that the abundance of Streptococcus was significantly reduced in PDAC, but not in CRC or RA (Figure 6E). In contrast, the abundance of Inocle-α was unchanged in the PDAC and RA groups compared to the HC group, but significantly reduced in the CRC group compared to the HC group (Figure 6E). These results suggest that the Inocle-α signature is particularly relevant to disorders occurring in the gastrointestinal tract.
[0240] <Correlation between Inocle-α and cancer - 2> Since there was a significant positive correlation (Spearman’s correlation = 0.14, p-value = 0.005) between the abundance of S. infantis (Relative abudance) and the abundance of Inocle-α (RPKM), S. infantis was included as a candidate host bacterium for Inocle-α. Therefore, the copy number of Inocle-α per copy of the S. infantis genome was calculated by calculating the ratio of S. infantis to Inocle-α.
[0241] Saliva samples were collected from head and neck cancer (HNC) patients including untreated patients, and metagenomic short reads were obtained. Next, 38 healthy controls (HC) with information on age, gender, and smoking status were compared with HNC patients to adjust for confounding factors. As a result, in the HNC group, the abundance of Inocle-α and the Inocle-α / S. infantis ratio were significantly decreased compared to the HC group, but the abundance of S. infantis was not significantly different between the HC group and the HNC group (Figure 17). It was also confirmed that the detection rate of Inocle-α was decreased in the HNC group (21%) compared to the HC group (60%) (Figure 17). Since cancer treatment may be a confounding factor, a comparison was made between the untreated HNC group and the HC group.
[0242] To analyze the association between Inocles and various diseases, the salivary metagenomic sequences of pancreatic ductal adenocarcinoma (PDAC), rheumatoid arthritis (RA), colorectal cancer patients (CRC), and healthy control groups (HC) in each study were downloaded from the NCBI database, and the Inocle-α / S. infantis ratio was calculated. As a result of the statistical test, it was shown that the abundance of S. infantis did not change significantly in all diseases. In contrast, in the HNC group (regardless of treatment) and the CRC group, Inocle-α / S. infantis was significantly decreased compared to the HC group (Figure 18).
[0243] [Example 7 Detection of Inocles] <Detection of Inocle-α by PCR> (Method) To evaluate primers capable of specifically detecting Inocle-α, 0.5 μL of a saliva sample containing Inocle-α (Inocle_004) as a template was added per reaction, and the detection of the target nucleic acid of Inocle-α was evaluated using a total of five primer sets shown in Table 2 below. Note that primer sets A to E were all designed not to anneal to the sequence on InoC but to anneal to regions other than InoC on Inocle-α.
[0244]
Table 2
[0245] KAPA HiFi TM Using the KAPA HiFi PCR Kit (Roche), a reaction solution containing the primer sets shown in combinations A to E of Table 2 (each primer at 0.4 μM) was prepared. Using an ABI / Pro Flex PCR system, the reaction solution was reacted under the following temperature cycle, and nucleic acid amplification was confirmed by electrophoresis. 95°C for 3 minutes, 35 cycles of 98°C for 20 seconds, 60°C for 30 seconds, and 72°C for 20 seconds
[0246] (Results) As a representative example, when nucleic acid amplification was performed using primer set B and the amplified product was detected by electrophoresis, the electrophoretic image obtained is shown in Figure 20. NTC is a negative control in which only the primer set was added without a saliva sample. Also, the amplified product was sequenced, and it was confirmed that the sequence of Inocle-α was amplified. In all cases using primer sets A, C, D, and E, amplified products were detected, similar to the case of primer set B.
[0247] (Design of Primer Sets and Probes for InoC Detection) A primer set and probe capable of specifically amplifying Inocle-α InoC were designed. More specifically, a pair of common portions was searched for in all of the nucleotide sequences represented by SEQ ID NOS: 30 to 45, one of which was designated as a forward primer and the other as a reverse primer. Furthermore, a portion of the nucleotide sequence sandwiched between the forward and reverse primers was used as a probe. The primer set was designed so that the Tm of the forward and reverse primers was approximately 55°C, and the probe was designed so that the Tm of the probe was approximately 58°C. The designed primer sets and probe sequences are shown in Table 3.
[0248] [Table 3]
[0249] The primer sets consisting of IC_a-1_F and IC_a-1_R, IC_a-2_F and IC_a-2_R, and IC_a-3_F and IC_a-3_R are primer sets that specifically recognize and amplify at least a portion of InoC of Inocle-α, i.e., a polynucleotide consisting of the base sequence represented by SEQ ID NOs: 30 to 45. Furthermore, IC_a-1_P, IC_a-2_P, and IC_a-3_P are probes capable of specifically binding to InoC of Inocle-α, that is, at least a part of the polynucleotide consisting of the nucleotide sequence represented by SEQ ID NOs: 30 to 45.
[0250] Herein, we identify and disclose a novel large ECE, a putative plasmid-like element, encoding multiple genes related to adaptation to intracellular and extracellular stresses in the oral cavity of many humans worldwide. We named this novel plasmid-like element "Inocle" and characterized it, focusing on its genetic and ecological significance. Furthermore, we uncovered significant associations between Inocles and human immune status and certain cancer types. These results suggest that human resident bacteria utilize large ECEs to adapt to changes in their habitat and human physiology.
[0251] Although Inocles can be detected in humans worldwide, they have not been identified or included in reference genomes due to the presence of ISs and hypermutated regions. The low scores of Inocles in plasmid prediction tools reflect the difficulty of detecting Inocles using current bioinformatics tools.
[0252] preNuc successfully reduced the amount of human DNA from saliva samples, contributing to high-throughput long-read metagenomic sequencing. The increased sequencing depth of inocles by preNuc treatment suggests that this method is effective for enriching intracellular ECEs.
[0253] The most notable feature of Inocles is its large contig size, up to 395 kb. This may make it one of the largest ECEs in human commensal bacteria. Based on its nucleotide composition, the inventors speculate that Inocles is a plasmid-like element and believe it can be classified as a megaplasmid. Megaplasmids have been discovered in several human pathogens and harbor multiple antimicrobial resistance genes, conferring multidrug resistance and enhancing infectivity. However, because Inocles does not contain any known antimicrobial or virulence genes, it is thought that Inocles may provide host bacteria with advantages distinct from those of previously identified human-associated megaplasmids.
[0254] The ORFs of Inocles were found to encode various stress response proteins. For example, RpoS is a sigma factor that globally regulates bacterial transcription and induces stress responses, promoting survival under environmental stress. In particular, a series of genes related to DNA damage repair pathways, including DNA polymerase III, gyrase, RecD, RecJ, and DinG, which are activated by oxidative stress, were identified. Furthermore, a peroxidase cofactor (NrdH) that reduces intracellular reactive oxygen species (ROS) levels was identified from Inocles-α, -β, and -γ. These results suggest that Inocles are important factors for resisting harsh environments such as oxidative stress and DNA damage in the oral environment (Figure 19).
[0255] Another feature of Inocle is that 20% of all genes are involved in cell wall / plasma membrane-related functions. The discovery of bifunctional transglycosylases, signal peptidases, sortase A, and LPXTG-like motifs suggests that the following processes occur on the plasma membrane: First, the bifunctional transglycosylase of Inocles promotes peptidoglycan biosynthesis in the host bacterium. After translation, LPXTG-like motif proteins are transported to the plasma membrane, where a signal peptidase cleaves the N-terminal signal peptide of the LPXTG-like motif proteins. Then, sortase A cleaves the LPXTG-like motif and anchors the mature proteins to peptidoglycan as cell surface proteins.
[0256] Several studies on human epithelial cell biofilms have revealed that sortase A contributes to bacterial adhesion and evasion of the human immune system. Although the functional role of LPXTG-like motif proteins remains to be elucidated, our experimental results suggest that Inocles cell wall-associated proteins are involved in Streptococcus colonization of human epithelial cells.
[0257] Some plasmids have functional redundancy with the host chromosome, contributing to increased genetic diversity and improved adaptive capacity. Homologous genes, such as DNA damage repair genes, oxidative stress-related genes, and cell wall-related genes, exist between the Inocles and Streptococcus chromosomes, and this functional redundancy is thought to contribute to improved adaptive capacity. On the other hand, most of the genes in Inocles have no homology to chromosomal genes in Streptococcus, suggesting that Inocles are functionally independent.
[0258] One of the highlights of this study is the elucidation of the relationship between Inocle-α and human physiological functions. Based on scRNA-seq and plasma proteome datasets, a clear positive correlation was observed between Inocle-α and signaling responses to infectious pathogens. While this study does not provide mechanistic insights into the interactions between Inocles, host bacteria, and the human immune system, these results suggest that Inocle-α is involved in latent or tolerant immune responses.
[0259] Another notable finding disclosed herein is the disease-specific reduction of Inocle-α. Specifically, a decrease in Inocle-α was observed in HNC and CRC, but not in PDAC and RA. One plausible explanation for these observations is that the immune status differs depending on the cancer type, and thus the requirement for Inocles changes as Streptococcus adapts to the immune response. Another possibility is that Inocles-positive individuals are less likely to develop the corresponding cancer because their immune system is already in an unfavorable state for cancer development. In either case, Inocles may serve as a biomarker for noninvasive testing of these cancers (Figure 19).
[0260] The findings provided herein are expected to pave the way for a deeper understanding of the multifaceted roles and biological significance of Inocles in regulating the adaptive capacity of oral bacteria, and ultimately their impact on human health.
[0261] The amino acid sequences deduced from the base sequences of the InoC genes of Inocle001 to 029 are shown in Table 4.
[0262] [Table 4-1]
[0263] [Table 4-2]
[0264] [Table 4-3]
[0265] The nucleotide sequences and SEQ ID NOs of the InoC genes of Inocle001 to 029 are shown in Table 5.
[0266] [Table 5-1]
[0267] [Table 5-2]
[0268] [Table 5-3]
[0269] [Table 5-4]
[0270] [Table 5-5]
[0271] [Table 5-6]
[0272] [Table 5-7]
[0273] [Table 5-8]
[0274] Representative amino acid sequences and SEQ ID NOs of functional genes of Inocles are shown in Table 6.
[0275] [Table 6-1]
[0276] [Table 6-2]
[0277] [Table 6-3]
[0278] [Table 6-4]
[0279] [Table 6-5]
[0280] [Table 6-6]
[0281] [Table 6-7]
[0282] [Table 6-8]
[0283] [Table 6-9]
[0284] The amino acid sequences of the protein family (PF) having an LPXTG-like motif are shown in SEQ ID NOs: 180 to 185, respectively, as shown in Table 7.
[0285] [Table 7]
[0286] As shown in Table 8, the base sequences of Inocle001 to 029 are represented by SEQ ID NOs: 149 to 177, respectively.
[0287] [Table 8]
[0288] [Table 9]
[0289] [Table 10]
[0290] [Table 11]
[0291] [Table 12]
[0292] Table 13
[0293] Table 14
[0294] Table 15-1
[0295] Table 15-2
[0296] Table 16
[0297] Table 17
Claims
1. A protein consisting of any one of the following amino acid sequences (a) to (c): (a) an amino acid sequence represented by any one of SEQ ID NOs: 1 to 29 (b) an amino acid sequence having an identity of 90% or more but less than 100% with an amino acid sequence represented by any one of SEQ ID NOs: 1 to 29; (c) an amino acid sequence represented by any one of SEQ ID NOs: 1 to 29 in which 1 to 10 amino acids are deleted, substituted, inserted or added;
2. A polynucleotide encoding the protein of claim 1.
3. The polynucleotide according to claim 2, consisting of a base sequence represented by any one of SEQ ID NOs: 30 to 58.
4. A circular DNA comprising the polynucleotide of claim 2.
5. The circular DNA according to claim 4, having a total length of 100 Kbp to 500 Kbp.
6. The circular DNA according to claim 5, comprising a polynucleotide encoding one or more proteins selected from the group of proteins consisting of the amino acid sequences represented by SEQ ID NOs: 61 to 148 or amino acid sequences having 50% or more identity to any of the amino acid sequences represented by SEQ ID NOs: 61 to 148.
7. The circular DNA according to claim 6, wherein the base sequence of the circular DNA is any one of the base sequences represented by SEQ ID NOs: 149 to 177.
8. A method for testing for head and neck cancer or colon cancer, comprising a step (C1) of measuring the amount of Inocles present in the saliva of a subject.
9. The testing method according to claim 8, wherein the Inocles are Inocles-α.
10. The testing method according to claim 8 or 9, wherein the step (C1) is a step of measuring the abundance of a part of the base sequence of Inocles.
11. The inspection method according to claim 10, wherein the step (C1) is a step of measuring the amount of InoC present.
12. The testing method according to claim 11, wherein the abundance of InoC is the abundance of one or more polynucleotides selected from a group of polynucleotides consisting of base sequences represented by SEQ ID NOs: 30 to 45.
13. The testing method according to claim 12, wherein the abundance of InoC is the total abundance of a group of polynucleotides consisting of base sequences represented by SEQ ID NOs: 30 to 45.
14. Furthermore, a step (C2) of measuring the amount of Streptococcus Infantis present in the saliva of the subject; and A step (C3) of calculating the amount of Inocles present per Streptococcus Infantis by dividing the amount of Inocles present obtained in the step (C1) by the amount of Streptococcus Infantis present obtained in the step (C2). The inspection method according to claim 8 , comprising:
15. The method for testing according to claim 8 or 14, wherein the amount of Inocles present in the step (C1) or the amount of Inocles present per Streptococcus infantis present in the step (C3) being smaller than a reference value indicates that the subject is highly likely to be suffering from head and neck cancer or colorectal cancer.
16. The testing method according to claim 15, wherein the reference value is the abundance of Inocles or the abundance of Inocles per Streptococcus infantis obtained from a healthy subject.
17. A saliva test kit for head and neck cancer or colon cancer testing, including an Inocles measurement reagent.
18. The saliva test kit according to claim 17, further comprising a reagent for measuring Streptococcus Infantis.
19. A marker for determining head and neck cancer or colorectal cancer, comprising the protein according to claim 1, the polynucleotide according to claim 2 or 3, or the circular DNA according to any one of claims 4 to 7.
20. A marker for determining whether a subject is affected with head and neck cancer or colorectal cancer, comprising the protein of claim 1, the polynucleotide of claim 2 or 3, or the circular DNA of any one of claims 4 to 7.
21. The protein according to claim 1, the polynucleotide according to claim 2 or 3, or the circular DNA according to any one of claims 4 to 7, which is used as a diagnostic marker for head and neck cancer or colon cancer.
22. A step (A1) of centrifuging an animal-derived specimen containing oral bacteria; and a step (A2) of treating the pellet obtained after removing the supernatant from the sample after the step (A1) with a nuclease; A method for producing a nucleic acid solution derived from oral bacteria, comprising:
23. The method of claim 22, wherein the sample is saliva.
24. The method of claim 22 or 23, wherein the animal is a human.
25. A primer set that specifically recognizes and amplifies at least a portion of a polynucleotide consisting of the base sequence represented by SEQ ID NOs: 149 to 177.
26. The primer set according to claim 25, wherein the primer set is selected from the following primer sets (A) to (E): (A) a polynucleotide consisting of the base sequence represented by SEQ ID NO: 190 and a polynucleotide consisting of the base sequence represented by SEQ ID NO: 191 (B) a polynucleotide consisting of the base sequence represented by SEQ ID NO: 192 and a polynucleotide consisting of the base sequence represented by SEQ ID NO: 193 (C) a polynucleotide consisting of the base sequence represented by SEQ ID NO: 194 and a polynucleotide consisting of the base sequence represented by SEQ ID NO: 195 (D) a polynucleotide consisting of the base sequence represented by SEQ ID NO: 196 and a polynucleotide consisting of the base sequence represented by SEQ ID NO: 197 (E) A polynucleotide consisting of the base sequence represented by SEQ ID NO: 198 and a polynucleotide consisting of the base sequence represented by SEQ ID NO:
199.
27. The primer set according to claim 25, wherein the primer set specifically recognizes and amplifies at least a part of a polynucleotide consisting of a base sequence represented by SEQ ID NOs: 30 to 58.
28. The primer set according to claim 27, wherein the primer set is selected from the following primer sets (Ia-A) to (Ia-D): Primer set (Ia-A): a polynucleotide consisting of the base sequence represented by SEQ ID NO: 200 and a polynucleotide consisting of the base sequence represented by SEQ ID NO: 201 Primer set (Ia-B): a polynucleotide consisting of the base sequence represented by SEQ ID NO: 202 and a polynucleotide consisting of the base sequence represented by SEQ ID NO: 203 Primer set (Ia-C): a polynucleotide consisting of the base sequence represented by SEQ ID NO: 205 and a polynucleotide consisting of the base sequence represented by SEQ ID NO: 206 Primer set (Ia-D): a polynucleotide consisting of the base sequence represented by SEQ ID NO: 208 and a polynucleotide consisting of the base sequence represented by SEQ ID NO: 209
29. A probe capable of specifically binding to at least a portion of a polynucleotide consisting of the base sequence represented by SEQ ID NOs: 149 to 177.
30. The probe according to claim 29, which is capable of specifically binding to at least a part of a polynucleotide consisting of a base sequence represented by any one of SEQ ID NOs: 30 to 58.
31. The primer set according to claim 30, wherein the probe is selected from the following probes (Ia-A) to (Ia-C): Probe (Ia-A): a polynucleotide consisting of the base sequence represented by SEQ ID NO: 204 Probe (Ia-B): Polynucleotide consisting of the base sequence represented by SEQ ID NO: 207 Probe (Ia-C): A polynucleotide consisting of the base sequence represented by SEQ ID NO: 210
32. A method for detecting Inocles, comprising a step of hybridizing the primer set according to any one of claims 25 to 28 or the probe according to any one of claims 29 to 31 with DNA contained in saliva.
33. A kit for detecting Inocles, comprising the primer set according to any one of claims 25 to 28 and / or the probe according to any one of claims 29 to 31.