Method for predicting onset and / or assessing risk of esophageal cancer, pharyngeal cancer, or oral cancer, and kit for predicting onset and / or assessing risk of esophageal cancer, pharyngeal cancer, or oral cancer

JPWO2025211164A5Pending Publication Date: 2026-04-14
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Filing Date
2025-03-18
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Current methods for detecting esophageal, pharyngeal, and oral cancer are invasive, costly, and lack accuracy due to unreliable self-reported lifestyle data and limited availability of blood samples for testing cancer driver mutations.

Method used

A non-invasive method using molecular barcode sequencing on buccal mucosa samples to identify mutations in specific genes and calculate allele frequencies, combined with genetic polymorphism analysis to assess cancer risk, achieving high accuracy through logistic regression and unique molecular identifiers.

Benefits of technology

The method provides objective and accurate risk assessment for esophageal, pharyngeal, and oral cancers with an AUC of 0.84 in ROC analysis, enabling early detection and prediction with minimal subject strain.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The present invention provides a method with which it is possible to objectively and accurately assess the risk of onset of esophageal cancer, pharyngeal cancer, and / or oral cancer, non-invasively and with little burden on the subject, and an onset prediction model having high accuracy for these cancers. The purpose of the present invention is also to provide a kit that can be used in the assessment method and the onset prediction model. The present invention is a method for predicting the onset and / or assessing the risk of esophageal cancer, pharyngeal cancer, or oral cancer in a subject, the method being characterized by comprising (A) a step for identifying a mutation in at least one gene selected from the group consisting of CHEK2, TP53, NOTCH 1, CDKN2A, NOTCH2, FAT1, PPM1D, AJUBA, ARID1A, ZFP36L2, ZNF750, and CUL3 in a nucleic acid specimen extracted from a biological sample derived from the oral mucosa of the subject and (B) a step for calculating and scoring the total (Sum of VAFs) allele frequency (VAF) and / or the total mutations (Number of mutations) of the mutation identified in step (A).
Need to check novelty before this filing date? Find Prior Art

Description

Method for predicting the onset and / or risk assessment of esophageal cancer, pharyngeal cancer, or oral cancer, and kit for predicting the onset and / or risk assessment of esophageal cancer, pharyngeal cancer, or oral cancer

[0001] The present invention relates to a method for predicting the onset of and / or assessing the risk of esophageal cancer, pharyngeal cancer, or oral cancer, and a kit for predicting the onset of and / or assessing the risk of esophageal cancer, pharyngeal cancer, or oral cancer.

[0002] Esophageal cancer is the sixth leading cause of cancer-related deaths worldwide, affecting approximately 604,100 people and resulting in 544,076 deaths each year (Non-Patent Document 1). Esophageal squamous cell carcinoma (ESCC) accounts for 85% of esophageal cancer cases, so there is a need for early detection of ESCC and accurate assessment of the risk of developing it.

[0003] Although endoscopic examination is effective for the early detection of esophageal cancer, problems have been pointed out, such as the invasive and painful nature of the examination, as well as the cost and time required (Non-Patent Document 2). Furthermore, lifestyle habits such as drinking and smoking, along with the presence of germline mutations in ALDH2 and ADH1B, which are important genes involved in alcohol metabolism, are known to be risk factors for ESCC (Non-Patent Documents 3-5). However, self-reported data on drinking and smoking are often inaccurate, which impairs the accuracy of risk assessment. Under these circumstances, establishing an objective and accurate assessment method to promote early detection of ESCC and assess the risk of developing the disease is an urgent issue.

[0004] In recent years, it has been discovered that clones carrying common cancer driver mutations expand in normal tissues (Non-Patent Documents 6-9). This clonal expansion correlates with known cancer risk factors such as aging, smoking, and alcohol consumption, suggesting that these driver mutation clones in normal tissues are likely to become the cells of origin for cancer (Non-Patent Documents 7 and 10). However, predicting cancer onset from the expansion of mutant clones in normal tissues remains difficult for many cancers, due in part to the limited availability of blood samples for testing (Non-Patent Documents 11 and 12).

[0005] CA Cancer J Clin 2021;71:209-49Gastrointest Endosc 2011;74:610-24 e2N Engl J Med 2003;349:2241-52Carcinogenesis 2017;38:859-72Cancer Lett 2009;275:240-6Nat Rev Cancer 2021;21:239-56Nature 2019;565:312-7Nature 2020;577:260-5Science 2020;370:75-82Nature 2020;578:266-72BMC Med 19, 243, doi:10.1186 / s12916-021-02109-y (2021) Clin Transl. Gastroenterol 9, 152, doi:10.1038 / s41424-018-0023-6 (2018)

[0006] In light of the above-mentioned circumstances, the present invention aims to provide a method that can objectively and accurately assess the risk of developing esophageal cancer, pharyngeal cancer, and / or oral cancer, which is non-invasive and places little strain on subjects, and a highly accurate prediction model for the onset of these cancers. Another object of the present invention is to provide a kit that can be used for the above-mentioned assessment method and onset prediction model.

[0007] To solve the above problems, the present inventors conducted various studies and found that clonal expansion in normal buccal mucosa can be used to objectively and accurately assess the risk of developing esophageal, pharyngeal, and oral cancer. The inventors hypothesized that, because the esophagus, pharynx, and oral cavity are covered with continuous squamous epithelium, similar genetic abnormalities might accumulate in normal esophagus and normal buccal mucosa. Experiments to confirm this assumption confirmed the expected similarity between the two. Furthermore, by performing molecular barcode sequencing on buccal mucosa samples collected noninvasively using swabs, they were able to detect minute mutant clones and successfully establish a method for objectively and accurately assessing the risk of developing not only oral cancer but also pharyngeal and esophageal cancer. The gist of the present invention is as follows.

[0008] [1] A method for predicting the onset and / or risk assessment of esophageal cancer, pharyngeal cancer, or oral cancer in a subject, comprising: (A) identifying mutations in at least one gene selected from the group consisting of CHEK2, TP53, NOTCH1, CDKN2A, NOTCH2, FAT1, PPM1D, AJUBA, ARID1A, ZFP36L2, ZNF750, and CUL3 in a nucleic acid sample extracted from a biological sample derived from the subject's oral mucosa; and (B) calculating and scoring the sum of allele frequencies (VAFs) of the mutations identified in step (A) and / or the total number of mutations. [2] The method of [1], wherein the identification of mutations in step (A) is performed by molecular barcode sequencing using unique molecular identifiers (UMIs). [3] The method of [1] or [2], wherein the scoring in step (B) is performed using the following formula: The cancer development risk score = a + ΣA(i) × Gene(i) Sum of variant allele frequencies (VAFs) or Total number of mutations (Number of mutations) (In the above formula, a, A(i) and Gene(i) are values ​​calculated by variable selection using Lasso and constant determination using logistic regression.) [4] The method according to [1] or [2], further comprising: (C) identifying genetic polymorphisms of aldehyde dehydrogenase ALDH2 and / or alcohol dehydrogenase ADH1B in the subject. [5] The method according to [1] or [2], wherein AUC in ROC analysis is 0.84 or more.[6] A kit used for predicting the onset and / or risk assessment of esophageal cancer, pharyngeal cancer, or oral cancer in a subject, comprising: (a) an instrument for collecting a biological sample from the oral mucosa of the subject; (b) a nucleic acid extraction reagent; (c) a molecular barcode sequencing reagent using a unique molecular identifier (UMI) to identify a mutation in at least one gene selected from the group consisting of CHEK2, TP53, NOTCH1, CDKN2A, NOTCH2, FAT1, PPM1D, AJUBA, ARID1A, ZFP36L2, ZNF750, and CUL3; and (d) software for analyzing the mutation. [7] A method for assisting in predicting the onset and / or risk assessment of esophageal cancer, pharyngeal cancer, or oral cancer in a subject, comprising: (A) identifying mutations in at least one gene selected from the group consisting of CHEK2, TP53, NOTCH1, CDKN2A, NOTCH2, FAT1, PPM1D, AJUBA, ARID1A, ZFP36L2, ZNF750, and CUL3 in a nucleic acid sample extracted from a biological sample derived from the subject's oral mucosa; and (B) calculating and scoring the sum of allele frequencies (VAFs) of the mutations identified in step (A) and / or the total number of mutations. [8] The method of [7], wherein the identification of mutations in step (A) is performed by molecular barcode sequencing using a unique molecular identifier (UMI). [9] The method of [7] or [8], wherein the scoring in step (B) is performed using the following formula: The cancer development risk score = a + ΣA(i) × Gene(i) Sum of allele frequencies (VAFs) of mutations (Sum of VAFs) or total number of mutations (Number of mutations) (In the above formula, a, A(i) and Gene(i) are values ​​obtained by variable selection using Lasso and constant determination using logistic regression.)

[10] (C) The method according to [7] or [8], further comprising a step of identifying genetic polymorphisms of aldehyde dehydrogenase ALDH2 and / or alcohol dehydrogenase ADH1B in the subject.

[11] The method according to [7] or [8], wherein the AUC in ROC analysis is 0.84 or more.

[0009] According to the determination method of the present invention, the risk of developing esophageal cancer, pharyngeal cancer and / or oral cancer can be objectively and accurately determined, and the onset of these cancers can be predicted with high accuracy.

[0010] Mutational profiles of esophageal driver genes in normal buccal mucosa. The numbers (top) and relative frequencies (bottom) of mutations assigned to SBS16 (alcohol profile), SBS4 (tobacco profile), SBS2 / 13 (APOBEC profile), and SBS1 / 5 (aging profile) are shown. The center shows the correlations between swab size, ESCC, HNSCC, germline risk mutations in ALDH2 and ADH1B, high-risk alcohol intake, high-risk smoking intensity, LVL grade C, and MCV ≥ 10. 6Information on fL is shown. Figure 1 shows the mean number and composition of variants assigned to each trait. (Germline Risk and Alcohol Consumption) Figure 2 shows the number of variants assigned to SBS16 (left) and SBS4 (right). Box plots represent the median, first quartile, and third quartile, with whiskers extending to the furthest value within the 1.5× interquartile range. P values ​​for significance between groups are also shown (Mann-Whitney U test). Figure 3 shows the mean number and composition of variants assigned to each trait. (Germline Risk and Smoking Intensity) Figure 4 shows the number of variants assigned to SBS16 (left) and SBS4 (right) plotted against alcohol consumption (ethanol units / week) and smoking intensity (pack-years), respectively, stratified by the presence or absence of germline risk variants. The p-values, coefficients, and intercepts of the linear mixed-effects model are shown for each risk group. Figure 5 shows the characteristics of positively selected genes in the buccal mucosa. Figure 6 shows the number of variants detected in each gene across all samples. Figure 7 shows the dN / dS values ​​for each gene and variant type. Asterisks indicate dN / dS > 1, q < 0.05. Calculations were performed using a two-sided likelihood ratio test and the Benjamini-Hochberg method for multiple testing correction using dndscv. This figure shows the distribution of mutations in the positively selected genes TP53, NOTCH1, PPM1D, and FAT1 in 387 buccal mucosa samples (top), 157 PNE samples (middle), and 519 ESCC samples (bottom). PNE refers to a physiologically normal esophagus. This figure shows the correlation between driver mutations in oral mucosa and risk factors for ESCC. The correlation between driver mutations in positively selected genes and known risk factors for ESCC was shown using a linear mixed-effects model for the number of driver mutations and the total VAF. Factors showing a significant correlation (p < 0.05) are shown. The number of driver mutations in TP53 and the total VAF are plotted against alcohol intake (ethanol units / week) in subjects with low-risk and high-risk ALDH2 variants. The p-values, coefficients, and intercepts of the linear mixed-effects models are shown. Correlation of driver mutations in positively selected genes with age and lifestyle risk factors for ESCC using linear mixed-effects models for the number of driver mutations and total VAF stratified by germline risk.Factors showing significant correlations (p<0.05) are shown. Figure 4 shows the correlation between buccal mucosal remodeling and lifestyle and clinical characteristics associated with ESCC. This figure plots the number of driver mutations against alcohol intake (left) and smoking intensity (right) for samples of subjects with low-risk and high-risk germline mutations. The p-values, coefficients, and intercepts of the linear mixed-effects model for each risk group are shown. Figure 4-A-1 shows the estimated coefficients for ethanol and tobacco in the model. 95% CIs are shown. Figure 4-B-1 shows the estimated coefficients for ethanol and tobacco in the model. 95% CIs are shown. Figure 4-B-1 shows histograms of the number of driver mutations (left) and the total VAF (right). ESCC is shown separately for subjects with and without germline risk mutations. Figure 5 shows the number of driver mutations (left) and the sum of VAFs (right) for subjects without ESCC, those with early-stage ESCC, or those with advanced ESCC, stratified by swab size. Figure 6 shows the number of mutations (left) and the sum of VAFs (right) for subjects without ESCC, those with ESCC only, or those with both ESCC and HNSCC. Box plots represent the median, first quartile, and third quartile, with whiskers extending to the furthest value within the 1.5× interquartile range. P values ​​for significant differences between samples are also shown (Mann-Whitney U test). Figure 5 shows the construction of an ESCC risk assessment model based on buccal mucosa driver mutation clones. This figure shows the proportion of trials in which variables were selected in a multivariate logistic regression model using LASSO for each swab type. The dotted line is drawn at a frequency of 0.6. Figure 6 shows a histogram of the c-index calculated from 1 million internal cross-validations for each swab type. The line indicates the median. ROC curves for each swab type. The threshold is the point farthest from the line, with a slope of 0.5. AUC, sensitivity, and specificity at the threshold are shown. Histograms of the sum of linear predictors are plotted for each swab type.Subjects without ESCC, subjects with ESCC only, and subjects with both ESCC and HNSCC were differentiated by color. Figure 6 shows the sample and sequencing results. A total of 396 samples were obtained from 110 subjects using three different swab sizes. After excluding samples with low consensus sequence depth (<1,800×) generated using molecular barcode sequencing, 387 samples were used for further analysis. Of these, 78 samples from 26 subjects were obtained from the opposite buccal mucosa of the same subject. Figure 6 shows the area of ​​the three different swab sizes. Calibration was performed. Figure 6 shows the sampling site. 20 mm. 2 , 28 mm 2 , 254 mm 2 Figure 1 shows a plot of DNA content for samples obtained using swabs. Figure 2 shows a bar graph showing raw depth (top), consensus depth, and final depth after downsampling for samples with consensus sequence depth ≥ 3,000x (bottom) by molecular barcode sequencing. Figure 3 shows a histogram of true-positive, false-negative, and false-positive mutations for simulated data using optimized filtering settings. Figure 4 shows the sensitivity and true-positive rate for mutations with a VAF ≥ 0.016. Figure 5 shows a histogram of VAF for mutations stratified by swab size. Figure 6 shows the mean depth and coverage for each WES sample. Figure 7 shows an overview of driver mutations in normal buccal mucosa (20mm swab size). 2 The clone size of mutations is plotted for each sample (top), and the colors indicate the driver genes for the three types of swabs. Heatmaps of the number (middle) and total (bottom) of VAFs for positively selected genes are shown. ESCC, HNSCC, risk mutations in ALDH2 and ADH1B, high-risk alcohol consumption, high-risk smoking intensity, MCV ≥ 10 6 The information on LVL grade C is also shown. Figure 1 shows an overview of driver mutations in normal buccal mucosa (swab size 28 mm). 2 ) Schematic of driver mutations in normal buccal mucosa (swab size 254 mm) 2Figure 8 shows the reproducibility of driver mutation clone detection in buccal mucosa using swabs. The number of driver mutations detected in paired samples is plotted for 10 positively selected genes and for driver mutations in NOTCH1, one of the three most frequently mutated genes, stratified by swab size. The line represents the regression line. The regression slope, intercept, and R2 value are shown. The total number of driver mutations detected in paired samples is plotted for 10 positively selected genes and for driver mutations in NOTCH1, one of the three most frequently mutated genes, stratified by swab size. The line represents the regression line. The regression slope, intercept, and R2 value are shown. The number of driver mutations detected in paired samples is plotted for 10 positively selected genes and for driver mutations in TP53 and CHEK2, one of the three most frequently mutated genes, stratified by swab size. The line represents the regression line. The regression slope, intercept, and R2 value are shown. The total number of driver mutations detected in paired samples is plotted for 10 positively selected genes and for driver mutations in TP53 and CHEK2, two of the three most frequently mutated genes, stratified by swab size. The line represents the regression line. The regression slope, intercept, and R2 value are shown. Figure 9 shows the distribution of mutations and negative selection in the buccal mucosa. This figure shows the distribution of NOTCH2, CHEK2, ARID1A, and CUL3 mutations in 387 buccal mucosa specimens (top), 157 PNE specimens (middle), and 519 ESCC specimens (bottom). PNE: normal esophagus. This figure shows the distribution of CDKN2A, AJUBA, and ZNF750 mutations in 387 buccal mucosa specimens (top), 157 PNE specimens (middle), and 519 ESCC specimens (bottom). The dN / dS values ​​for each gene and mutation type are shown. Asterisks indicate dN / dS<1 and q<0.05. The correlation coefficients were calculated using a two-sided likelihood ratio test with dndscv and a Benjamini-Hochberg multiple testing correction. Figure 10 shows the correlation between driver mutations and alcohol consumption. The figure shows the number and total VAFs of FAT1 mutations plotted against alcohol consumption (units of ethanol / week) in a sample of subjects with low-risk and high-risk ALDH2 mutations. The p-values, coefficients, and intercepts of the linear mixed-effects model are shown.This figure shows the number and total VAFs of PPM1D mutations plotted against alcohol consumption (ethanol units / week) in samples of subjects with low-risk and high-risk ALDH2 mutations. The p-values, coefficients, and intercepts of the linear mixed-effects model are shown. Figure 11 shows the relationship between buccal mucosal remodeling and ESCC characteristics. This figure plots the number of driver mutations (left) and total VAFs (right) for single specimens stratified by germline risk. P-values ​​for significance between groups are shown (Mann-Whitney U test). This figure shows the number (left) and total VAFs (right) plotted according to MCV values ​​(MCV<10). 6 fL vs MCV≧10 6fL). This figure shows the number (left) and total (right) of VAFs plotted according to LVL grade (grade A / B vs. grade C). Box plots represent the median, first quartile, and third quartile, with whiskers extending to the furthest value within the 1.5× interquartile range. p-values ​​for significant differences between samples are also shown (Mann-Whitney U test). Figure 12 shows the construction of an ESCC risk assessment model based on buccal mucosa driver mutation clones. This figure shows the proportion of studies in which variables were selected in the multivariate logistic regression model using LASSO for each swab type. The dotted line represents a frequency of 0.8. For each swab type, receiver operating characteristic curves (ROC) curves for the discovery and additional cohorts were constructed using models constructed based on the matched data from the discovery cohort. The points indicate the threshold determined as the point closest to (0, 1) on the ROC curve for the discovery cohort. The AUC, sensitivity, and specificity at the threshold are shown. For each swab type, ROC curves were plotted for subjects with and without germline risk mutations using the same model as in Figure 12-B. AUC, sensitivity, and specificity at the same threshold are shown. Histograms of the sum of 890 linear predictors in the discovery cohort were plotted by swab type, distinguishing between subjects without ESCC, subjects with ESCC only, and subjects with both ESCC and HNSCC. Histograms of the c-index calculated from 10,000 internal cross-validations for each swab type are shown. The dotted line indicates the median. Sensitivity (left) and specificity (right) for the discovery and expansion cohorts are shown using a model constructed based on the discovery cohort data, and propensity score-matched samples were evaluated for both cohorts. Histograms of the sum of linear predictors in the expansion cohort were plotted by swab type, distinguishing between subjects without ESCC, subjects with ESCC only, and subjects with both ESCC and HNSCC.

[0011] The present invention will be described in detail below. The terms used in this specification are to be interpreted in the sense commonly used in the art unless otherwise specified.

[0012] The present invention relates to a method for predicting the onset and / or risk assessment of esophageal, pharyngeal, or oral cancer in a subject, comprising: (A) identifying a mutation in at least one gene selected from the group consisting of CHEK2, TP53, NOTCH1, CDKN2A, NOTCH2, FAT1, PPM1D, AJUBA, ARID1A, ZFP36L2, ZNF750, and CUL3 in a nucleic acid sample extracted from a biological sample derived from the oral mucosa of the subject; and (B) calculating and scoring the sum of allele frequencies (VAFs) and / or the total number of mutations (Number of mutations) identified in step (A). The present invention may further comprise (C) identifying genetic polymorphisms of aldehyde dehydrogenase ALDH2 and / or alcohol dehydrogenase ADH1B in the subject.

[0013] The inventors have found that similar genetic abnormalities accumulate in the esophagus, pharynx, and oral cavity, which are covered with continuous squamous epithelium, due to the same environmental influences brought about by saliva. Therefore, they have developed a method of the present invention that enables the prediction and early detection of not only oral cancer but also esophageal cancer and pharyngeal cancer by examining genetic abnormalities in oral mucosa, such as the buccal mucosa. According to the method of the present invention, molecular barcode sequencing is performed using a biological sample derived from oral mucosa noninvasively collected from a subject using a swab or other method, allowing the detection of minute mutant clones and the objective and accurate assessment of the risk of developing not only oral cancer but also pharyngeal cancer and esophageal cancer. The method of the present invention is described in detail below.

[0014] [Step (A)] This step is a step of identifying a mutation in at least one gene selected from the group consisting of CHEK2, TP53, NOTCH1, CDKN2A, NOTCH2, FAT1, PPM1D, AJUBA, ARID1A, ZFP36L2, ZNF750, and CUL3 in a nucleic acid specimen extracted from a biological sample derived from the oral mucosa of the subject.

[0015] In the present invention, a biological sample derived from oral mucosa refers to a sample containing cells and tissues derived from the buccal mucosa, tongue, floor of the mouth, gums, hard palate, etc. Among these, buccal mucosa is preferred from the viewpoint of the stability of data obtained in the method of the present invention.

[0016] The method for collecting the biological sample is not particularly limited, and for example, the sample can be collected using a wiping swab (wiping device). Here, a swab is a wiping device consisting of a wiping portion made of a core material, fiber, resin, etc., and the shape is not limited, but examples include cotton swabs and sponge brushes. The term "swab" may also refer to a biological sample collected using a swab as a wiping device. When collecting the biological sample, the amount, quality, etc. of the specimen to be collected may be adjusted depending on the size (area) of the wiping portion of the swab or the like to be used. The size (area) is 1 mm 2 ~1000mm 2 5 mm 2 ~500mm 2 Preferably, it is 10 mm 2 ~400mm 2 More preferably, it is 15 mm 2 ~300mm 2 More preferably, it is 20 mm 2 ~260mm 2 It is even more preferable that 2 ~100mm 2 It is particularly preferable that 2 ~30mm 2 When collecting the biological sample, the wiping device such as a swab is rotated in the same area of ​​the oral cavity a certain number of times, for example, 2 to 20 times, 5 to 15 times, 8 to 12 times, or 10 times.

[0017] In the present invention, a nucleic acid sample refers to a sample containing nucleic acid extracted and separated from the above-mentioned biological sample. Here, nucleic acid refers to a molecule in which nucleotides, each consisting of a base, sugar, and phosphate, are linked by phosphodiester bonds, and includes ribonucleic acid (RNA) and deoxyribonucleic acid (DNA). The method for extracting and separating nucleic acid from the above-mentioned biological sample is not particularly limited, and any method known to those skilled in the art can be used. Commercially available kits (e.g., QIAGEN DNA Micro Kit) may also be used.

[0018] In this step, mutations in at least one gene selected from the group consisting of CHEK2, TP53, NOTCH1, CDKN2A, NOTCH2, FAT1, PPM1D, AJUBA, ARID1A, ZFP36L2, ZNF750, and CUL3 are identified in the nucleic acid sample. CHEK2 (Checkpoint kinase 2) is a gene involved in DNA damage response and encodes a protein that regulates cell cycle arrest and apoptosis. Mutations in the CHEK2 gene are known to increase the risk of breast cancer and ovarian cancer. TP53 is a gene encoding p53, a protein that plays an important role in cellular DNA damage response. Mutations in the TP53 gene are known to be involved in many types of cancer. NOTCH1 is part of the Notch signaling pathway and is involved in cell differentiation, growth, and cell death. Mutations and activation of NOTCH1 have been reported in certain types of leukemia and solid tumors. CDKN2A is a gene involved in cell cycle regulation and encodes two proteins, p16INK4a and p14ARF. Mutations in the CDKN2A gene are known to increase the risk of melanoma and other cancers. Like NOTCH1, NOTCH2 is part of the Notch signaling pathway and is involved in cell fate determination. Mutations in NOTCH2 are associated with certain cancers and diseases, such as head and neck cancer and skin cancer. FAT1 is a gene encoding a protein involved in cell adhesion. Mutations in FAT1 have been reported to be associated with head and neck cancer and other cancers. PPM1D (Protein phosphatase, Mg2+ / Mn2+ dependent, 1D) is a gene encoding a protein involved in cellular and stress responses. Mutations in the PPM1D gene are associated with several cancers, including breast and ovarian cancer. AJUBA is a gene encoding a protein involved in cytoskeleton regulation and signal transduction pathways. Mutations in the AJUBA gene may be associated with the progression of certain cancers, such as head and neck cancer. ARID1A (AT-rich interactive domain-containing protein 1A) is a gene encoding a protein involved in chromatin remodeling.Mutations in ARID1A have been found in several cancers, including ovarian and gastric cancer. ZFP36L2 (ZFP36 ring finger protein-like 2) is a gene encoding an RNA-binding protein with a CCCH zinc finger structure that binds to the 3' untranslated region of mRNAs containing AU-rich elements (AREs) and promotes the degradation of target mRNAs. ZNF750 is a gene encoding a protein involved in cell differentiation. Mutations in the ZNF750 gene are particularly associated with skin cancer. CUL3 (Cullin 3) is a gene encoding a protein involved in protein degradation, part of a ubiquitin ligase complex. Mutations in the CUL3 gene are known to be associated with the risk of hypertension and certain cancers, including lung and kidney cancer.

[0019] The NCBI (National Center for Biotechnology Information) Reference Sequence accession numbers for the above genes are as follows: CDKN2A; NM_000077 TP53; NM_000546 ARID1A; NM_006015 CUL3; NM_001257197 PPM1D; NM_003620 FAT1; NM_005245 ZFP36L2; NM_006887 ZNF750; NM_024702 AJUBA; NM_032876 NOTCH2; NM_024408 NOTCH1; NM_017617 CHEK2; NM_001005735

[0020] The method for identifying mutations in each of the genes in the nucleic acid sample is not particularly limited, but from the perspective of superior sensitivity, molecular barcode sequencing using unique molecular identifiers (UMIs) or targeted sequencing targeting esophageal driver mutations is preferred, with molecular barcode sequencing using unique molecular identifiers (UMIs) being particularly preferred. These methods identify mutations in CHEK2, TP53, NOTCH1, CDKN2A, NOTCH2, FAT1, PPM1D, AJUBA, ARID1A, ZFP36L2, ZNF750, and CUL3 in a subject's oral mucosal sample, and calculate the total number of driver mutations (Number of drivers) and the sum of allele frequencies (VAFs) of driver mutations (Sum of VAFs). In the next step [B], the degree of clonal expansion can be scored and assessed.

[0021] In the method of the present invention, it is preferable to identify mutations in at least one gene selected from the group consisting of CHEK2, TP53, NOTCH1, NOTCH2, FAT1, PPM1D, and CUL3 among the above-mentioned 12 genes, and as described below, depending on the swab size, it is preferable to identify mutations in PPM1D, NOTCH1, CUL3, and TP53, identify mutations in NOTCH2, FAT1, CUL3, CHEK2 and TP53, or identify mutations in TP53 and CHEK2.

[0022] A library is prepared using the nucleic acid sample (DNA). After fragmenting the DNA, strand-specific molecular indexing can be performed by ligating adapters with double-stranded 8-bp unique molecular indexes (UMIs) to independently tag the upper and lower strands of the inserted DNA. The indexed library can be hybridized, captured, and sequenced using xGen™ predesigned gene capture pools (IDTs) designed for the gene. Sequencing reads can be aligned to the human reference genome (GRCh37) using a program such as BWA-MEM (v.0.7.17) to identify mutations. This method allows for highly sensitive mutation reads.

[0023] The gold standard is a somatic mutation with a variant allele frequency (VAF) of ≥ 0.1 detected in a sufficient number of ESCC samples. To simulate low-VAF mutations, ESCC sample data and normal sample data are mixed at a ratio of 0.2–1%, and mutations are called using the virtual data. Filtering parameters are determined to improve sensitivity and reduce the false-positive rate. The final parameters can be set to the following values: (i) VAF ≥ 0.0016 and VAF ≤ 0.3, (ii) the maximum number of mutant reads detected in 20 controls generated from blood DNA ≤ 2, (iii) the number of reads resequenced to artifact sequences inferred from cruciform DNA formation ≤ 2, (iv) the number of reads resequenced to the original sequence ≥ 2, and (v) a VAF significantly higher than expected for errors (p ≤ 10-2.5) determined using the EBCall algorithm based on the empirical VAF distribution from 20 controls.

[0024] Validation analysis may also be performed using digital droplet PCR (ddPCR) using a QX200 (Bio-Rad) or similar device. ddPCR analysis can be performed by randomly selecting mutations in any of the above genes and using sample DNA and control DNA samples. The threshold for distinguishing mutant allele signals from wild-type allele signals is determined based on the signal distribution obtained from the control sample. For example, the following values ​​can be set as validation criteria: (i) number of mutant signals in the swab sample ≥ 3; (ii) number of mutant signals in the control ≤ 1; (iii) VAF in the control swab sample ≥ 5 × VAF. Assuming no copy number change, the mutant clone size can be calculated by multiplying the swab area by 2 × VAF for mutations on the autosome and X chromosome (female subjects) and by VAF for mutations on the X chromosome (male subjects).

[0025] [Step (B)] This step is a step of calculating and scoring the sum of allele frequencies (VAFs) of mutations identified in the above step (A) and / or the total number of mutations. Since the number and size of gene mutations detected differ depending on the swab size, the type and number of genes used in the calculation formula for scoring differ for each swab size. The risk of developing esophageal cancer, pharyngeal cancer, or oral cancer is estimated based on the 12 genes positively selected in the buccal mucosa, which are 20 mm or less. 2 For swab: 4 genes (PPM1D, NOTCH1, CUL3, TP53) 28 mm 2 For swab: 5 genes (NOTCH2, FAT1, CUL3, CHEK2, TP53) 254 mm 2 In the case of swabs, it is proportional to the number of mutations or the mutant cell fraction (sumVaf) of two genes (TP53, CHEK2). Therefore, the calculation formula is the sum of the sum of the allele frequencies (VAFs) of mutations in each gene (Sum of VAFs) and the total number of mutations (Number of mutations) multiplied by a constant, as shown below.

[0026] Risk of developing the above cancer = a + ΣA(i) × Gene(i) Sum of allele frequencies (VAFs) or Sum of mutations

[0027] Here, a, A(i), and Gene(i) can be obtained by variable selection using Lasso (narrowing down from 12 genes) (Journal of the Royal Statistical Society, Series B (Methodological), Vol. 58, No. 1 (1996), pp. 267-288) and constant determination using logistic regression.

[0028] For people whose presence or absence of esophageal, pharyngeal, or oral cancer is known, genome analysis of oral mucosa such as buccal mucosa can be performed, and the number of genes and percentage of cells mutated in the cancer can be determined by variable selection using Lasso to create the logistic regression formula. This can be used to identify gene mutations in oral mucosa such as buccal mucosa of people whose presence or absence of pharyngeal or oral cancer is unknown, and the formula can be used to calculate a score and assess the risk of developing the cancer.

[0029] Specifically, the score can be calculated based on the following formula: 2 When using swab: -1.76765654788505 + 39.2897138502103 × PPM1D_driver_sumVaf (total number of allele frequencies of PPM1D mutations) + 0.0353564566438162 × NOTCH1_driver_count (total number of NOTCH1 mutations) + 1.47997141135261 × CUL3_driver_count (total number of CUL3 mutations) + 0.193911099674833 × TP53_driver_count (total number of TP53 mutations) ・28mm 2 Using swab: -2.27530895775624 + 0.366458763289863 × NOTCH2_driver_count (total number of NOTCH2 mutations) + 0.0934156421436881 × FAT1_driver_count (total number of FAT1 mutations) + 1.19979494864608 × CUL3_driver_count (total number of CUL3 mutations) + 0.273419476668119 × CHEK2_driver_count (total number of CHEK2 mutations) + 0.255985620159922 × TP53_driver_count (total number of TP53 mutations) ・254 mm 2Using swab: -1.16032619761279 + 0.182876161159443 × TP53_driver_count (total number of TP53 mutations) + 0.345307999692175 × CHEK2_driver_count (total number of CHEK2 mutations)

[0030] Regarding the risk of developing esophageal cancer, pharyngeal cancer, or oral cancer, if the score obtained based on the above formula is less than -2, the risk is determined to be low (low risk group); if it is -2 or more but less than 3, the risk is determined to be medium (moderate risk group); and if it is 3 or more, the risk is determined to be high (high risk group). In the low risk group, the subject has not developed the cancer at the time of testing and is determined to have a low risk of developing it in the future. In the medium risk group, the subject is determined to have a high risk of developing the cancer in the future even if they have not developed the cancer at the time of testing. In the high risk group, the subject is determined to have a high possibility of developing the cancer at the time of testing, or to have an extremely high risk of developing the cancer in the future even if they have not developed the cancer.

[0031] To increase versatility, an additional cohort of 112 people was analyzed in the same way, and the resulting data was used to predict the presence or absence of esophageal cancer with high accuracy in both the discovery cohort and the additional cohort. The more versatile formula is as follows:

[0032] 20mm 2 For swab: 3 genes (NOTCH1, CUL3, TP53) 28 mm 2 For swab: 2 genes (FAT1, TP53) 254 mm 2 For swab: 2 genes (TP53, CHEK2)

[0033] ・20mm 2 Using swab: -2.24942 + 0.08366 × NOTCH1_driver_count (total number of NOTCH1 mutations) + 1.90984 × CUL3_driver_count (total number of CUL3 mutations) + 0.11889 × TP53_driver_count (total number of TP53 mutations) ・28 mm 2When using swab: -1.8594 + 0.2233 × FAT1_driver_count (total number of FAT1 mutations) + 0.2647 × TP53_driver_count (total number of TP53 mutations) ・254 mm 2 With swab: -1.31835 + 0.32676 × CHEK2_driver_count (total number of CHEK2 mutations) + 0.14762 × TP53_driver_count (total number of TP53 mutations)

[0034] Regarding the risk of developing esophageal cancer, pharyngeal cancer, or oral cancer, if the score obtained based on the above formula is less than 0, the risk is determined to be low (low risk group); if it is 0 or more but less than 2, the risk is determined to be medium (moderate risk group); and if it is 2 or more, the risk is determined to be high (high risk group). In the low risk group, the subject has not developed the cancer at the time of testing and is determined to have a low risk of developing it in the future. In the medium risk group, the subject is determined to have a high risk of developing the cancer in the future even if they have not developed the cancer at the time of testing. In the high risk group, the subject is determined to have a high possibility of developing the cancer at the time of testing, or to have an extremely high risk of developing the cancer in the future even if they have not developed the cancer.

[0035] [Step (C)] This step is a step of identifying genetic polymorphisms of aldehyde dehydrogenase ALDH2 and / or alcohol dehydrogenase ADH1B in a subject. By combining the values ​​related to driver mutations obtained in the above steps [A] and [B] and the calculated score with information on genetic polymorphisms of aldehyde dehydrogenase ALDH2 and / or alcohol dehydrogenase ADH1B in the subject obtained in this step, the accuracy of determining the risk of developing esophageal cancer, pharyngeal cancer, or oral cancer can be further improved.

[0036] The method for identifying genetic polymorphisms of aldehyde dehydrogenase ALDH2 and alcohol dehydrogenase ADH1B is not particularly limited, and any conventionally known method can be used.

[0037] A large body of prior art literature has supported the finding that genetic polymorphisms in the alcohol dehydrogenase (ADH1B) and aldehyde dehydrogenase (ALDH2) genes increase the risk of cancer in alcohol users, particularly in Asian populations. Humans carrying the ADH1B*1 / *1 or ALDH2*1 / *2 genotypes are at increased risk of upper respiratory tract, gastrointestinal, and other cancers when consuming alcohol (Druesne-Pecollo N. et al., Lancet Oncol. 2009, 10(2), 173-80). Single nucleotide polymorphisms (SNPs) in ADH1B and ALDH2 play important roles in alcohol metabolism. The ADH1B*1 allele, located on chromosome 4q22, slows down ethanol metabolism, while the ADH1B*2 allele accelerates it. The alleles located on chromosome 12q24 are ALDH2*1, which allows normal metabolism of acetaldehyde, and ALDH2*2, which causes defective metabolism of acetaldehyde.

[0038] In addition to the above steps, information on alcohol consumption, smoking amount, and the presence or absence of a flushing reaction due to alcohol can be obtained through interviews or questionnaires and used to determine the risk of developing esophageal cancer, pharyngeal cancer, or oral cancer.

[0039] The reliability of the method of the present invention for assessing the risk of developing esophageal cancer, pharyngeal cancer, or oral cancer can be demonstrated by ROC analysis, which creates an ROC curve. According to the method of the present invention, the AUC in the ROC analysis is 0.84 or higher, as shown in the Examples, demonstrating that the method is highly reliable. An AUC of 0.9 to 1.0 is considered to be "very high" reliability, 0.8 to 0.9 is considered to be "high" reliability, 0.7 to 0.8 is considered to be "medium" reliability, and 0.5 to 0.7 is considered to be "low" reliability.

[0040] "AUC" is an abbreviation for "area under a curve." AUC refers to the area under a receiver operating characteristic (ROC) curve. An ROC curve is a plot of the true positive rate against the false positive rate for different possible cut points of a diagnostic test. The ROC curve shows the trade-off between sensitivity and specificity, which depends on the cut point chosen (any increase in sensitivity may be accompanied by a decrease in specificity). AUC is a measure of the accuracy of a diagnostic test, with higher values ​​indicating higher accuracy, with a maximum of 1 and 0.5 for a random test.

[0041] In the method of the present invention, mutations in at least one gene selected from the group consisting of CHEK2, TP53, NOTCH1, CDKN2A, NOTCH2, FAT1, PPM1D, AJUBA, ARID1A, ZFP36L2, ZNF750, and CUL3 are identified in the oral mucosa of a subject by the above-mentioned steps [A] and [B]. The sum of allele frequencies (VAFs) of the identified mutations (sum of VAFs) and / or the total number of mutations (number of mutations) are calculated and scored, and then compared with the average reference value for healthy individuals to objectively and accurately assess the risk of developing not only oral cancer but also pharyngeal cancer and esophageal cancer (oral mucosal clonal expansion model). Furthermore, accuracy can be further improved by using a multifactor model (genetic polymorphism + oral mucosal clonal expansion model, oral mucosal clonal expansion model + oral mucosal clonal expansion model, or genetic polymorphism + oral mucosal clonal expansion model) that combines information on ALDH2 and ADH1B genetic polymorphisms, alcohol intake, smoking amount, and flushing response information obtained through a medical interview.

[0042] According to the determination method of the present invention, by using a biological sample derived from the oral mucosa, samples can be collected non-invasively and easily. Furthermore, by identifying and quantifying mutations in at least one gene selected from the group consisting of CHEK2, TP53, NOTCH1, CDKN2A, NOTCH2, FAT1, PPM1D, AJUBA, ARID1A, ZFP36L2, ZNF750, and CUL3, the risk of developing not only oral cancer but also pharyngeal cancer and esophageal cancer can be objectively and accurately determined, making it possible to predict the onset of these cancers with a high degree of accuracy.

[0043] <Method for assisting in predicting the onset and / or risk assessment of esophageal cancer, pharyngeal cancer, or oral cancer> The above-mentioned method for predicting the onset and / or risk assessment of esophageal cancer, pharyngeal cancer, or oral cancer does not include a diagnostic method by a physician, but to make this more clear, it can also be described as a method for assisting in predicting the onset and / or risk assessment of esophageal cancer, pharyngeal cancer, or oral cancer. The substantial content of the method is the same as the above-mentioned method for predicting the onset and / or risk assessment of esophageal cancer, pharyngeal cancer, or oral cancer. For specific explanations, the content of the section on the method for predicting the onset and / or risk assessment of esophageal cancer, pharyngeal cancer, or oral cancer can be applied as is.

[0044] <Kit for predicting the onset of and / or assessing the risk of esophageal cancer, pharyngeal cancer, or oral cancer> The present invention also includes a kit used for predicting the onset of and / or assessing the risk of esophageal cancer, pharyngeal cancer, or oral cancer in a subject, comprising: (a) an instrument for collecting a biological sample from the oral mucosa of the subject; (b) a nucleic acid extraction reagent; (c) a molecular barcode sequencing reagent using a unique molecular identifier (UMI) for identifying a mutation in at least one gene selected from the group consisting of CHEK2, TP53, NOTCH1, CDKN2A, NOTCH2, FAT1, PPM1D, AJUBA, ARID1A, ZFP36L2, ZNF750, and CUL3; and (d) software for analyzing the mutation.

[0045] (a) An instrument for collecting a biological sample derived from the oral mucosa of the subject. An example of an instrument for collecting the biological sample is a wiping swab (wiping instrument). Here, a swab is a wiping instrument consisting of a wiping part made of a core material, fiber, resin, etc., and the shape is not limited, but examples include cotton swabs and sponge brushes. When collecting the biological sample, the amount, quality, etc. of the specimen to be collected can be adjusted by changing the size (area) of the wiping part of the swab or the like to be used. The size (area) is 1 mm 2 ~1000mm 2 5 mm 2 ~500mm 2 Preferably, it is 10 mm 2 ~400mm 2 More preferably, it is 15 mm 2 ~300mm 2 More preferably, it is 20 mm 2 ~260mm 2 It is even more preferable that 2 ~100mm 2 It is particularly preferable that 2 ~30mm 2 It is most preferable that:

[0046] (b) Nucleic Acid Extraction Reagents As nucleic acid extraction reagents, commercially available reagents conventionally known for use in nucleic acid extraction and separation can be used.

[0047] (c) A commercially available reagent can be used as a molecular barcode sequencing reagent using a unique molecular identifier (UMI) to identify mutations in at least one gene selected from the group consisting of CHEK2, TP53, NOTCH1, CDKN2A, NOTCH2, FAT1, PPM1D, AJUBA, ARID1A, ZFP36L2, ZNF750, and CUL3. Mutations in TP53, CHEK2, NOTCH1, NOTCH2, FAT1, and AJUBA include all nonsynonymous substitutions, while mutations in PPM1D, ARID1A, ZFP36L2, CKDN2A, and ZNF750 are only truncating mutations, and mutations in CUL3 are only missense mutations.

[0048] (d) The software for analyzing the above mutations is not particularly limited, but a preferred example is The Genomon Pipeline (v.2.6.3) (http: / / github.com / Genomon-Project).

[0049] The present invention will be described in more detail below based on examples, but the present invention is not limited to these examples.

[0050] 1. Confirmation of the Similarity between the Esophagus and the Pharynx (1) Comparison of Macroscopic Images: mLVL. In normal esophagus, glycogen production results in intense staining with iodine dispersion. However, cancerous areas exhibit iodine-unstained lesions due to reduced glycogen production. This technique is also used clinically in the diagnosis and treatment of esophageal cancer. Interestingly, in patients with esophageal cancer, the background esophagus exhibits multiple lugol voiding lesions (mLVL), which are small, iodine-unstained lesions. Endoscopically, typical cases of esophageal cancer with 10 or more LVL lesions per image are classified as LVL grade C, cases with zero LVL lesions per image are classified as LVL grade A, and all other lesions are classified as grade B. In 2022, Toranomon Hospital reported a 55% concordance rate between the esophageal and pharyngeal LVL grades (Esophagus 2022; 19: 460-68). In our own cases at Kyoto University Hospital, the concordance rate when classified as LVL grade C or grade A / B was high at 87% (27 / 31). These results demonstrate the similarity between the esophagus and pharynx in terms of LVL grade.

[0051] (2) Microscopic Comparison: Genetic Abnormalities in LVL As mentioned above, typical cases of esophageal cancer present with mLVL. It has been reported in Japan and China that the esophageal LVL that constitutes mLVL is composed of esophageal driver mutations, primarily TP53, a driver mutation in esophageal cancer (Cell Rep Med 2023; 4; 101-68, Cancer Medicine 2021; 10:3545-3555). In our own case data, we performed whole-exome analysis of esophageal and pharyngeal LVL and compared the driver mutations. The number of mutations in both esophageal and pharyngeal LVL was found to be consistent between the esophagus and pharynx, with the highest number of mutations in TP53, followed by NOTCH1 and FAT1 (unpublished data). These findings demonstrate the similarities between the esophagus and pharynx in terms of genetic abnormalities.

[0052] 2. Establishment of a method for assessing esophageal cancer risk using normal buccal mucosa (1) Materials and Methods (Patient Samples) Sixty-two subjects with ESCC or a history of ESCC were enrolled at Kyoto University Hospital and Kobe Asahi Hospital between December 2019 and December 2022 (Figure 6-A). Among them, 35 subjects had current or previous head and neck squamous cell carcinoma (HNSCC). An additional 48 subjects without ESCC confirmed by endoscopy within 12 months of sample collection were enrolled during the same period (Figure 6-A). Because ESCC patients frequently engage in heavy drinking, patients with relatively low alcohol intake were prioritized for inclusion. In addition, subjects without ESCC were included to achieve a uniform age distribution and a similar proportion of heavy drinkers. Informed consent was obtained from all participants in accordance with procedures approved by the Kyoto University Internal Review Board (G0645).

[0053] Furthermore, as an additional cohort, 59 ESCC patients and 53 ESCC-free subjects from Kyoto University Hospital and Kobe Asahi Hospital were analyzed between January 2023 and April 2024.

[0054] Buccal mucosa samples were taken at 20, 28 and 254 mm 2Three sizes of swabs with an area of ​​10 mm (Fig. 6B, C) were used for collection. To collect samples, the swab was rotated 10 times over the same area of ​​the buccal mucosa. Each swab sampled a different area. White blood cells were collected from 20 subjects as controls. Blood tests were performed within 6 months of sampling in all but two cases (9 and 20 months). Lugol's staining images of the esophagus were obtained during endoscopy within 12 months of sampling in all subjects with ESCC and six subjects without ESCC. The grade of iodine-unstained Lugol voiding lesions (LVLs) was independently assessed by three experienced endoscopists based on defined criteria (Grade A: no LVLs, Grade B: 1–9 LVLs per endoscopic finding, Grade C: ≥10 LVLs). High-risk drinking habits (heavy drinking) were defined as an average alcohol intake of ≥18 units / week (1 unit = 22 g of ethanol), and high-risk smoking habits (heavy smoking) were defined as a smoking intake of ≥30 pack-years. One pack-year corresponds to one pack per day for one year. ESCC patients were staged according to the WHO Classification of Tumors of the Digestive System (5th edition). Stages 0 and 1 were classified as early ESCC, and stages 2-4 were classified as advanced ESCC.

[0055] In our previous study, we analyzed 519 ESCC and 157 physiologically normal esophageal samples. We selected the following 26 genes that were frequently mutated (≥2%) and significantly mutated via MutSig CV (q < 0.05) or dndscv (q < 0.1) (Table 1). ALDH2 and ADH1B were also analyzed.

[0056]

[0057] In addition to the genes listed in Table 1 above, analysis of an additional cohort newly demonstrated positive selection of ZFP36L2 in the buccal mucosa.

[0058] DNA was extracted using a QIAGEN DNA Micro Kit (Figure 6-D), followed by library preparation using 25-50 ng of DNA. After fragmentation with an E220evolution Focused-ultrasonicator (Covaris), strand-specific molecular indexing was performed by ligating adapters with double-stranded 8-bp unique molecular indexes (UMIs) to independently tag the upper and lower strands of the insert DNA (xGen™ Prism DNA Library Prep Kit, IDT). The indexed library was hybridized and captured using an xGen™ predesigned gene capture pool (IDT) designed for 28 genes. Sequencing was performed on a DNBSEQ-G400RS (MGI Tech) in 150-bp paired-end mode. The target raw depth for sequencing was 20,000×, with a median raw depth of 40,671× (Figure 6E). Sequencing reads were aligned to the human reference genome (GRCh37) using BWA-MEM (v. 0.7.17). Read families containing top and bottom strands derived from the same original molecule were determined based on the start-stop positions of the reads and unique UMI sets (double sequencing) using fgbio (v. 1.4.0) and picard (v. 2.24.0). Next, consensus reads were generated, correcting for artifacts, only for read families with two or more top reads and one or more bottom reads; the rest were discarded. The script used to construct the consensus reads has been deposited in Zenodo (doi: 10.5281 / zenodo.10023627). Samples with a consensus read depth of 1,800x or greater (110 specimens, 387 cases) were used in this study (Figure 6A). Because mutation detection sensitivity depends on the read depth, samples with a read depth of 3,000x or greater were randomly downsampled to 2,500x, resulting in a median consensus read depth of 2,494x (Figure 6E).

[0059] Somatic mutations were called using the Genomon pipeline (v. 2.6.3) (http: / / github.com / Genomon-Project). Modifications were made to EBFilter (https: / / github.com / kwatanabe-ku) to accurately process high-depth data (v. 0.2.5). Modifications were made to GenomonMutationFilter to evaluate artifacts associated with cruciform DNA formation (v. 0.4.0). To determine filtering parameters, 12 ESCC samples and 16 independent normal blood samples were sequenced in the same manner as the buccal swab samples. A somatic mutation with a variant allele frequency (VAF) ≥ 0.1 detected in the 12 ESCC samples was used as the gold standard. To simulate low VAF mutations, ESCC sample data and normal sample data were mixed at a ratio of 0.2-1%. Mutations were called in the virtual data, and filtering parameters were determined to improve sensitivity and reduce the false-positive rate. The final parameters were: (i) VAF ≥ 0.0016 and VAF ≤ 0.3, (ii) maximum number of mutant reads detected in 20 controls generated from blood DNA ≤ 2, (iii) number of reads resequenced to artifact sequences inferred from cruciform DNA formation ≤ 2, (iv) number of reads resequenced to the original sequence ≥ 2, and (v) VAF significantly higher than expected for errors (p ≤ 10), as determined using the EBCall algorithm based on the empirical VAF distribution from 20 controls. -2.5 This filtering process yielded a sensitivity of 0.914 and a true positive rate of 0.964 (Figure 6F). To remove artifacts, we further analyzed mutation loci that were mutated in six or more samples (n = 318). Loci in which 70% or more of the mutations were located within 50 bases from either end of the read were excluded (n = 17). Considering the variant reads across all mutant samples, variants with 8 or more supporting reads and a strand ratio of 0 or 1 were also considered artifacts and were excluded (n = 27).

[0060] Validation analysis was also performed using digital droplet PCR (ddPCR) with a QX200 (Bio-Rad). The inventors selected TP53 mutations using validated probes available from the manufacturer. 124 mutations were randomly selected. ddPCR analysis was performed using 50 ng of sample DNA and three control DNA samples. The threshold for distinguishing mutant allele signals from wild-type allele signals was determined based on the signal distribution obtained from the control sample. Validation criteria were as follows: (i) number of mutant signals in the swab sample ≥ 3; (ii) number of mutant signals in the three control samples ≤ 1; and (iii) VAF in the swab samples of the three control samples ≥ 5 × VAF. The true positive rate was 91.1%. Mutant clone size was calculated by multiplying the swab area by 2 × VAF for mutations on the autosomes and X chromosome (female subjects) and VAF for mutations on the X chromosome (male subjects), assuming no copy number changes.

[0061] Mutational signature analysis was performed using the R package MutationalPatterns (v. 3.10.0). Somatic mutations detected in samples with one or more single-nucleotide mutations were decomposed into COSMIC SBS1, SBS5, SBS2, SBS13, SBS4, and SBS16, which are known signatures in ESCC and physiologically normal esophagus (6, 32, 33). Single samples were used for comparisons between germline and lifestyle risk groups, while all samples were used for LME model analysis.

[0062] (Whole Exome Sequencing (WES) analysis) Samples were taken from 0.2 mm of the buccal mucosa. 2Punch biopsies (KAI BP-A05F) were collected, followed by DNA extraction using a DNA micro kit. Matched blood samples were then collected. After enzymatic fragmentation of 1.2-10 ng of DNA, libraries were prepared using the xGen™ Lotus DNA Library Prep Kit (IDT). Indexed libraries were hybridized and captured using the xGen™ Exome Hyb Panel v2 (IDT). Sequencing was performed in 150 bp paired-end mode on a DNBSEQ-G400RS. The target sequencing depth was 100x, with an actual average depth of 85x (21-139x) (Figure 6H). Mutation calling was performed using the Genomon pipeline as previously described. The filtering parameters used in Genomon were as follows: (i) mapping quality ≥ 20; (ii) base call quality ≥ 15; (iv) total reads ≥ 10 and variant reads ≥ 3; (v) VAF ≥ 0.05 (tissue cases) and ≤ 0.02 (normal control cases); (vi) strand ratio ≠ 0 or 1; (vii) EBcall p ≤ 10-4 using WES data of unpaired peripheral blood samples (n = 16); and (viii) Fisher test p ≤ 10-1.

[0063] To identify expanded clones across multiple samples, each mutation detected in a sample was examined in all remaining samples. If a mutation detected in one or more samples by Genomon was also present in another sample within the biopsy sample (VAF ≥ 0.05), the mutation was considered a true positive. The mutant cell fraction was calculated by multiplying the sample area by 2 × VAF for mutations on the autosomes and X chromosome (female subjects) and by VAF for mutations on the X chromosome (male subjects), since there were no copy number changes assessed by CNACS.

[0064] (External dataset) Somatic mutations identified in 519 ESCC samples and 157 physiologically normal esophageal samples were obtained from our previous study (Nature 2019;565:312-7). MutSig CV and dndscv results were also obtained.

[0065] (Germline risk loci for ALDH2 and ADH1B) ALDH2 rs671 and ADH1B rs1229984 are common functional variants in East Asian populations. Risk and protective genotypes were defined as rs671 (AG, risk; AA and GG, protective) and rs1229984 (GG, risk; AA and AG, protective), respectively. ALDH2 rs671 (AG) and ADH1B rs1229984 (GG) were defined as high-risk germline variants (Gastroenterology 2009;137:1768-75).

[0066] (Model Construction for ESCC Risk Assessment) We first performed univariate analysis to determine the presence or absence of ESCC using a logistic regression model, taking into account factors including germline risk variants in ALDH2 and ADH1B and the number and sum of VAFs for driver variants in each positively selected gene. Significant factors (p<0.05) were selected for further analysis, and 75% of the samples were randomly selected. Multivariate analysis was performed using the logistic regression model. Next, the optimal combination of factors was selected using the least absolute shrinkage and adaptive selection operator (LASSO) method using the R package glmnet (v.4.1.8). A "germline" model (gene polymorphism) was constructed using stable combinations of factors selected in 60% or more of 1 million tests. Internal cross-validation was performed to assess the validity of the model by randomly dividing all samples into training (75%) and validation (25%) sets 1 million times. A test model was constructed using the training set, and the c-index of the validation set was calculated. A final model was constructed using all isolated samples of each swab type, and the optimal threshold was determined using the ROC curve, where the point furthest from a line with a slope of 0.5 was selected to maximize sensitivity. The R code used to construct the model has been deposited at Zenodo (doi: 10.5281 / zenodo.10576948).

[0067] As a more versatile model, a model was created using factors selected in 80% or more of the tests using the LASSO method.

[0068] Statistical analysis: Statistical analysis was performed using R (v. 4.2.3). Unless otherwise specified, all P values ​​were calculated using two-sided analysis. Mann-Whitney U test or chi-square test was used for group comparison. LVL grade comparison was performed excluding missing values ​​(complete case analysis). To examine the homogeneity of buccal mucosal remodeling, we applied linear regression to compare the number and total variant allele frequencies (VAFs) from the left and right buccal mucosa. The LME model was evaluated using the R package lmerTest (v. 3.1.3). All samples were included in the analysis, with subject and swab size as random effects.

[0069] (2) Results (Somatic Mutations in ESCC Driver Genes in the Buccal Mucosa) A total of 110 patients were enrolled, including 62 with ESCC and 48 without ESCC (Figure 6-A). To analyze the influence of lifestyle factors on buccal mucosa remodeling, the two groups of patients were matched for age and heavy drinking habits. No significant differences were found in age, gender, heavy drinking habits, or heavy smoking habits. Compared with subjects without ESCC, the ESCC group had a higher frequency of germline risk mutations in ALDH2 (86% vs. 27%, p<0.01) and ADH1B (21% vs. 4%, p=0.01). Among subjects who reported an alcohol flushing response, a surrogate marker for low-type ALDH2 alleles, 9% (6 of 66) had only wild-type ALDH2 alleles, whereas 20% (8 of 44) of non-responders had low-type ALDH2 alleles. Because the size of the mutant clones was unknown, three swab sizes (20, 28, and 254 mm²) were used to obtain buccal mucosa samples (Figure 6B). For each subject, samples were obtained from different parts of the central region of the buccal mucosa, while three additional samples were obtained from a subset of patients (n = 26) using three swab sizes from the other side (Figure 6C). DNA was extracted from the swabs (Figure 6D), followed by library construction using unique molecular identifiers (UMIs). Target capture sequencing was performed at a raw target depth of 20,000 (Figure 6E) for 26 driver genes positively selected in ESCC and physiologically normal esophagus, as well as ALDH2 and ADH1B. Consensus reads were then generated using both the top and bottom strands, each tagged with a unique set of UMIs. Samples (n = 387 from 110 subjects) with a consensus read depth of 1,800x or greater were included for further analysis, while samples with a depth of 3,000x or greater were downsampled to 2,500x to prevent increased sensitivity for mutation detection with low VAF. Filtering settings for mutation calling were determined using simulated data generated by combining ESCC and germline data generated in the same way at ratios ranging from 0.2 to 1.0%.The sensitivity for mutations with a VAF of ≥ 0.0016 (Figure 6-F) was 0.91, and the true positive rate was 0.96. Digital droplet PCR experiments verified that the VAF range of 113 of the 124 TP53 mutations was 0.0016–0.0108 (91.1%).

[0070] Somatic mutations were detected in all samples (median, 25 (1-133)). A total of 12,865 mutations were detected, including 11,816 nonsynonymous mutations and 1,049 synonymous mutations, with a VAF range of 0.0016 to 0.12 (median, 0.0025). The somatic mutations detected in each sample were resolved into COSMIC single-base substitution (SBS) signatures (31): SBS1 / 5, SBS2 / 13, SBS4, and SBS16 (Figure 1A). These signatures have been identified in ESCC or physiologically normal esophagus and are associated with aging, APOBEC enzyme activity, smoking, and alcohol consumption, respectively. Among subjects with germline risk mutations, heavy drinkers had a significantly higher number of SBS16 mutations than light drinkers (p = 0.003), but this difference was not observed among subjects without mutations (p = 0.51) (Figure 1B, C). Heavy smokers tended to have a higher number of SBS4 mutations than nonsmokers, both among subjects with and without germline risk mutations (Fig. 1C, D). A linear mixed-effects model (LME) was applied to the number of SBS16 mutations in each group, stratified by the presence or absence of germline risk mutations, with alcohol consumption as a fixed effect and subject and swab size as random effects (Fig. 1E). Among subjects with germline risk mutations, alcohol consumption was positively correlated with the number of SBS16 mutations (p = 0.018), but not among subjects without mutations (p = 0.20). We also analyzed the relationship between the number of SBS4 mutations and smoking intensity. The number of SBS4 mutations increased with smoking intensity regardless of germline mutation (p = 0.11, high-risk germline subjects; p = 0.009, low-risk germline subjects), but the increase was more significant in subjects with germline risk mutations (Fig. 1E).

[0071] (Clonal expansion driven by ESCC driver mutation in buccal mucosa) 254 mm 2 Among the clones detected using swabs, the largest clone was 48.38 mm2 90% of cases have a nonsense mutation in PPM1D and are 3 mm in size. 2 In contrast, the results were 20 mm or less (Figs. 6G and 7). 2 and 28 mm 2 The largest clones detected using swabs were 4.74 mm 2 and 4.16 mm 2 and most clones were 0.34 mm in size. 2 (median, 0.10 mm 2 and 0.14 mm 2 ) (Figs. 6-G, 7), the lower limit of detectable VAF (≥ 0.0016 or 0.81 mm 2 ) for 254mm 2 The clone size detected by swabbing was not detectable. To verify the clone size detected by swabbing, a 58-year-old patient who underwent buccal mucosal resection for buccal mucosal cancer was enrolled. Normal mucosa was observed without cancer. After separating the epithelium from the stroma, a punch biopsy was performed to measure 0.2 mm 2 Samples (n = 40) were collected in a 1 mm grid and subsequently subjected to WES (Fig. 6H). Clones with TP53 mutations were observed in four samples, with most mutations confined to individual samples. The median size of the mutant clones was 0.1 mm. 2 (VAF 0.25), which is 28 mm 2 and 20 mm 2 These results were comparable to those detected using swabs. 2 It contains many clones with sizes of 10 mm and some clones are sometimes tens of mm. 2 Therefore, it shows that the size of the 2 From 254 mm 2 The observed area of ​​the mutant clones was sufficiently large compared to that of the mutant clones, providing a comprehensive view of buccal mucosal remodeling. To evaluate the reproducibility of our method, we applied linear regression to compare the number and total number of VAFs detected for mutations between samples obtained from the left and right buccal mucosa of the same subject, which showed a high level of concordance (Fig. 8).

[0072] NOTCH1 was the gene with the highest mutation frequency, followed by TP53 and CHEK2, with rates of 38.3%, 28.6%, and 5.9%, respectively (Figure 2A). To examine the mode of natural selection in the buccal mucosa, we performed dN / dS analysis using the R package dndscv. Of the 26 genes analyzed, six genes, TP53, CHEK2, NOTCH1, NOTCH2, FAT1, and AJUBA, showed dN / dS > 1 (q < 0.1) for both missense and protein-truncating mutations, indicating positive selection of mutations in these genes (Figure 2B; Table 1). The mutation distribution of each gene largely overlapped between buccal mucosa, physiologically normal esophagus, and ESCC samples (Figure 2C; Figure 9A). On the other hand, six genes, PPM1D, ARID1A, CDKN2A, ZFP36L2, ZNF750, and CUL3, showed positive selection against either missense or truncating mutations. In contrast, nine genes, PPM1D, ZFP36L2, ARID1A, KDM6A, PTEN, PLXNB2, NOTCH3, PIK3CA, and PTCH1, showed evidence of negative selection (dN / dS < 1, q < 0.1) against missense and / or truncating mutations (Figure 9B). Notably, the dN / dS values ​​for PPM1D and ARID1A were < 1 for missense mutations and > 1 for truncating mutations, suggesting different effects of mutation types on natural selection for the two genes. We focused our analysis on driver mutations, defined as positively selected mutations, in 12 genes (nonsynonymous mutations in six genes; truncating mutations in PPM1D, ARID1A, CDKN2A, ZFP36L2, and ZNF750; and missense mutations in CUL3).

[0073] (Relationship between buccal mucosa mutations and ESCC risk factors) We analyzed the correlation of known risk factors for ESCC, age, alcohol consumption, smoking intensity, and germline mutations in ALDH2 and ADH1B, with driver mutations in each positively selected gene in terms of the number of driver mutations and their sum of VAFs, using a linear mixed-effects model with random effects for subject and swab size (Figure 3A). Among risk factors, germline risk mutations in ALDH2 showed the most positive correlations, with correlations with eight genes. In addition to germline risk mutations in ALDH2, driver mutations in TP53, FAT1, and PPM1D also showed positive correlations with alcohol consumption. The effect of alcohol consumption on the number and sum of VAFs for mutations in each of the three genes was analyzed, stratified by the presence or absence of germline risk mutations in ALDH2 (Figure 3B, Figures 11A and 11B). Positive correlations were observed only among subjects with risk mutations. These results suggest that alcohol consumption and germline ALDH2 risk variants combine to drive positive selection for variants in these genes. In contrast, among subjects without germline risk variants, driver variants were not correlated with alcohol consumption but were correlated with smoking intensity (Figure 3C). Across risk factors, age was not significantly correlated with driver variants, likely due to the subjects' advanced age (>40 years).

[0074] Correlation of buccal mucosal remodeling with lifestyle and ESCC characteristics. To quantify buccal mucosal remodeling due to positively selected clones, the total number of driver mutations and the sum of their VAFs were calculated for each sample. Compared with subjects without germline risk mutations, subjects with germline mutations showed greater remodeling (p<0.05) (Figure 11-A). Next, we examined the relationship between buccal mucosal remodeling and ESCC. Subjects with extensive remodeling were more likely to have ESCC than subjects with less extensive remodeling (Figure 4C). We analyzed the relationship between driver mutations in the buccal mucosa and clinical features associated with ESCC. Compared with subjects without ESCC, patients with early or advanced ESCC had significantly higher numbers and sums of driver mutation VAFs (p<0.05) (Figure 4D). There was no significant difference between early and advanced ESCC, suggesting that even in patients with early ESCC, the buccal mucosa underwent extensive remodeling due to driver mutation clones. Elevated mean corpuscular volume (MCV) has been found to correlate with the presence of ESCC. MCV≧10 6 Subjects with MCV <10 6 The number and total number of driver mutation VAFs were higher in subjects with ESCC than in subjects with ESCC (Figure 4B). Head and neck squamous cell carcinoma (HNSCC) often occurs synchronously or metachronously in ESCC patients. Subjects with both ESCC and HNSCC had significantly higher number and total driver mutation VAFs than subjects with ESCC only (p<0.05), indicating highly remodeled regions in subjects with bifocal cancer (Figure 4E). LVL grade C is associated with a higher risk of metachronous ESCC and HNSCC than LVL grade A / B. The number and total number of driver mutation VAFs tended to be higher in grade C individuals than in grade A / B individuals, but the difference was not statistically significant (Figure 4C).

[0075] Buccal mucosal remodeling as a marker of ESCC risk. Buccal mucosal remodeling caused by positively selected clones was associated with lifestyle and germline risk factors, the presence of ESCC, and ESCC-related clinical characteristics, motivating the establishment of a predictive model for ESCC using genetic swab data. Given that the contribution of each risk factor to mutant clones differed among positively selected genes (Figure 3A), we first attempted to identify a set of factors reflecting ESCC risk. Univariate logistic regression was applied to individual samples to assess the presence or absence of ESCC by considering the number and total number of variable impact factors (VAFs) of driver mutations in each gene. Significant variables from the univariate analysis were retained. Next, a multivariate logistic regression model was applied to the presence or absence of ESCC in 75% of randomly selected subjects. Variable selection was performed using least absolute shrinkage and adaptive selection operator (LASSO) methods to prevent model overfitting.

[0076] Variables selected in 60% or more of the 1 million tests were used to build a predictive model (Figure 5A). 2 The total number of mutations in TP53, CUL3, and NOTCH1, and the total VAF of PPM1D mutations were selected from the swabs. 2 The swab showed TP53, CHEK2, CUL3, FAT1, and NOTCH2 mutations, 254 mm 2 The swabs were selected for the number of CHEK2 and TP53 mutations. To evaluate model performance, 75% of the subjects were randomly selected to train a "genetic" model based on the selected variables, and the concordance index (c-index) was calculated using the remaining 25% of the subjects. This procedure was repeated 1 million times, resulting in a high c-index (Figure 5B).

[0077] A predictive model was also created using variables selected in 80% or more of the cases as a more versatile model that could be validated in additional cohorts (Figure 12A).

[0078] Finally, ROC curves were constructed for the three swab types for all subjects (Fig. 5C). To maximize sensitivity, the threshold was determined as the point on the ROC curve farthest from the line with a slope of 0.5. 2The swab size showed the highest area under the curve (AUC) value of 0.92 and sensitivity of 0.95 among the three swab sizes, suggesting that this size is optimal for obtaining a comprehensive landscape of buccal mucosal remodeling. Because the range of buccal mucosal remodeling was more extensive in patients with both ESCC and HNSCC (Figure 4E), these patients had more highly correlated values ​​of the linear predictor (Figure 5D), whereas subjects with lower values ​​showed a lower proportion of ESCC. These results indicate that the model accurately reflects ESCC risk. Furthermore, the AUC of the highly generalizable model was 0.84–0.85 (Figures 12B–D).

[0079] The scores in Figure 5D were calculated (scored) using the following formula, which was created according to a logistic regression model based on the clonal expansion (total number of mutations and allele frequencies) of 12 genes present in the buccal mucosa. The values ​​were divided into bins and the frequency distribution is shown in Figure 5D. Specifically, the following formula was used for each swab size in the scoring.

[0080] ・20mm 2 swab: -1.76765654788505 + 39.2897138502103 × PPM1D_driver_sumVaf (total number of allele frequencies of PPM1D mutations) + 0.0353564566438162 × NOTCH1_driver_count (total number of PPM1D mutations) + 1.47997141135261 × CUL3_driver_count (total number of CUL3 mutations) + 0.193911099674833 × TP53_driver_count (total number of TP53 mutations) ・28mm 2swab: -2.27530895775624 + 0.366458763289863 × NOTCH2_driver_count (total number of NOTCH2 mutations) + 0.0934156421436881 × FAT1_driver_count (total number of FAT1 mutations) + 1.19979494864608 × CUL3_driver_count (total number of CUL3 mutations) + 0.273419476668119 × CHEK2_driver_count (total number of CHEK2 mutations) + 0.255985620159922 × TP53_driver_count (total number of TP53 mutations) ・254 mm 2 swab: -1.16032619761279 + 0.182876161159443 × TP53_driver_count ((total number of TP53 mutations)) + 0.345307999692175 × CHEK2_driver_count (total number of CHEK2 mutations)

[0081] Regarding the risk of developing esophageal cancer, pharyngeal cancer, or oral cancer, if the score obtained based on the above formula is less than -2, the risk is determined to be low (low risk group); if it is -2 or more but less than 3, the risk is determined to be medium (moderate risk group); and if it is 3 or more, the risk is determined to be high (high risk group). In the low risk group, the subject has not developed the cancer at the time of testing and is determined to have a low risk of developing it in the future. In the medium risk group, the subject is determined to have a high risk of developing the cancer in the future even if they have not developed the cancer at the time of testing. In the high risk group, the subject is determined to have a high possibility of developing the cancer at the time of testing, or to have an extremely high risk of developing the cancer in the future even if they have not developed the cancer.

[0082] To increase versatility, an additional cohort of 112 people was analyzed in the same way, and the resulting data was used to predict the presence or absence of esophageal cancer with high accuracy in both the discovery cohort and the additional cohort. The more versatile formula is as follows: 2 For swab: 3 genes (NOTCH1, CUL3, TP53) 28 mm 2For swab: 2 genes (FAT1, TP53) 254 mm 2 For swab: 2 genes (TP53, CHEK2)

[0083] ・20mm 2 Using swab: -2.24942 + 0.08366 × NOTCH1_driver_count (total number of NOTCH1 mutations) + 1.90984 × CUL3_driver_count (total number of CUL3 mutations) + 0.11889 × TP53_driver_count (total number of TP53 mutations) ・28 mm 2 When using swab: -1.8594 + 0.2233 × FAT1_driver_count (total number of FAT1 mutations) + 0.2647 × TP53_driver_count (total number of TP53 mutations) ・254 mm 2 With swab: -1.31835 + 0.32676 × CHEK2_driver_count (total number of CHEK2 mutations) + 0.14762 × TP53_driver_count (total number of TP53 mutations)

[0084] Regarding the risk of developing esophageal cancer, pharyngeal cancer, or oral cancer, if the score obtained based on the above formula is less than 0, the risk is determined to be low (low risk group); if it is 0 or more but less than 2, the risk is determined to be medium (moderate risk group); and if it is 2 or more, the risk is determined to be high (high risk group). In the low risk group, the subject has not developed the cancer at the time of testing and is determined to have a low risk of developing it in the future. In the medium risk group, the subject is determined to have a high risk of developing the cancer in the future even if they have not developed the cancer at the time of testing. In the high risk group, the subject is determined to have a high possibility of developing the cancer at the time of testing, or to have an extremely high risk of developing the cancer in the future even if they have not developed the cancer.

[0085] Figure 13-A shows a histogram of the c-index calculated from 10,000 internal cross-validations for each swab type. Figure 13-B also shows the results of evaluating the sensitivity and specificity of the discovery and expansion cohorts for propensity-matched samples using a model developed based on the discovery cohort data. Figure 13-C also shows a histogram of the sum of linear predictors in the expansion cohort, plotted by swab type, distinguishing between subjects without ESCC, subjects with ESCC only, and subjects with both ESCC and HNSCC.

[0086] (3) Discussion In this study, 2 From 254 mm 2 We developed a noninvasive method to obtain a tissue remodeling landscape by using high-sensitivity molecular barcode sequencing of buccal mucosa samples collected using a range of swabs. We showed that buccal mucosa remodeling by mutant clones reflects a range of lifestyle risk factors, including alcohol consumption and smoking, as well as germline risk mutations in ALDH2 and ADH1B. These factors are major risk factors for cancer in solid organs such as the esophagus, lung, and bladder, suggesting the usefulness of buccal mucosa remodeling as a surrogate marker for assessing cancer risk. Indeed, this model based on buccal mucosa remodeling accurately predicted the presence of ESCC. To our knowledge, this is the first study to use clonal expansion in normal tissue to noninvasively assess cancer risk in solid organs.

[0087] The buccal mucosa is composed of numerous small clones, allowing the detection of a sufficient number of somatic mutations for dN / dS analysis. Of the 26 genes positively selected in ESCC or physiologically normal esophagus, 11 genes were positively selected and 9 genes were negatively selected in the buccal mucosa. This suggests that, as previously known, there are common and distinct selection mechanisms between cancer and normal tissues. Among the positively selected genes, TP53, CHEK2, and PPM1D were incorporated into a model for assessing ESCC risk. These three genes belong to the p53 signaling pathway, suggesting that functional loss of this pathway may play an important role in carcinogenesis. Tissue remodeling by mutant clones is accelerated by aging and exposure to tobacco, alcohol, and other substances. The effect of alcohol consumption on the number and size of SBS16 mutation and driver mutation clones was observed only in subjects with germline risk mutations, suggesting that germline and environmental factors jointly influence tissue remodeling and providing insight into the mechanisms behind the recently reported observation that BRCA1 / 2 germline risk mutations increase gastric cancer risk only in chronic Helicobacter pylori infection (N Engl J Med 388, 1181-1190 (2023)). Furthermore, alcohol exposure did not promote clonal expansion in subjects without germline risk mutations, consistent with the low prevalence of ESCC in countries with low incidence of germline risk mutations. Risk stratification of ESCC based on lifestyle and alcohol flushing response is currently in use. However, accuracy is limited, in part due to difficulties in accurately quantifying drinking and smoking habits and estimating germline risk from traits such as alcohol flushing response. TP30 mutations have been reported to be more frequently detected in the esophageal epithelium of patients with higher LVL grades. However, due to the invasiveness of endoscopy, mutation analysis of the esophageal mucosa has not been utilized for ESCC risk assessment. However, optical properties of the buccal mucosa measured using reflectance spectroscopy have been reported to reflect cancerous areas (Clin Transl Gastroenterol 9, 152 (2018)). It is also a non-invasive method for ESCC risk assessment.However, although no direct comparison has been performed, the reported sensitivity was lower than that of the present method. Recently, cell-free DNA methylation analysis has been used to detect ESCC (BMC Med 19, 243 (2021)), but its sensitivity for early-stage ESCC was low due to the low proportion of cancer-derived cell-free DNA. Because this study was a case-control study, it was difficult to perfectly match cases and controls, which may lead to selection bias. Indeed, the proportion of subjects with germline risk mutations significantly differed between ESCC patients and non-ESCC patients. However, our model, which does not use germline information, demonstrated high accuracy in predicting ESCC. To establish the usefulness of buccal mucosal remodeling, future large-scale cohort studies are needed to track the development of ESCC, similar to other smoking- and alcohol-related cancers.

[0088] Early detection and prevention of cancer are essential for future improvement. This study suggests that measuring buccal mucosal remodeling may allow for early intervention to identify high-risk individuals even before the onset of ESCC, reduce exposure to lifestyle risk factors, and facilitate endoscopic surveillance.

Claims

1. A method for predicting the onset and / or assessing the risk of esophageal cancer, pharyngeal cancer, or oral cancer in a subject, (A) A step of identifying mutations in at least one gene selected from the group consisting of CHEK2, TP53, NOTCH1, CDKN2A, NOTCH2, FAT1, PPM1D, AJUBA, ARID1A, ZFP36L2, ZNF750, and CUL3 in a nucleic acid sample extracted from a biological sample derived from the oral mucosa of the subject. (B) A step of calculating and scoring the sum of allele frequencies (VAFs) and / or the number of mutations of the mutations identified in step (A), and (C) Step of determining the LVL grade of the esophagus of the subject. A method characterized by including

2. (A) The method according to claim 1, wherein the identification of mutations in step (A) is performed by molecular barcode sequencing using a unique molecular identifier (UMI).

3. (B) The method according to claim 1 or 2, wherein scoring in step (B) is performed by the following calculation formula. The risk score for the above cancer is calculated as follows: a + ΣA(i) × Gene(i) - Sum of VAFs or Number of mutations (In the above formula, a, A(i), and Gene(i) are values ​​obtained by variable selection using Lasso and constant determination using logistic regression.)

4. (D) A step of identifying gene polymorphisms of aldehyde dehydrogenase ALDH2 and / or alcohol dehydrogenase ADH1B in the subject. The method according to claim 1 or 2, further comprising:

5. The method according to claim 1 or 2, wherein the AUC in the ROC analysis is 0.84 or higher.

6. A kit used for predicting the onset and / or assessing the risk of esophageal cancer, pharyngeal cancer, or oral cancer in subjects, (a) an instrument for collecting a biological sample derived from the oral mucosa of the subject, (b) Nucleic acid extraction reagents (c) A reagent for molecular barcode sequencing using a unique molecular identifier (UMI) to identify mutations in at least one gene selected from the group consisting of CHEK2, TP53, NOTCH1, CDKN2A, NOTCH2, FAT1, PPM1D, AJUBA, ARID1A, ZFP36L2, ZNF750, and CUL3. (d) Software for analyzing the above mutations, (e) Reagents for determining the LVL grade in the esophagus of the subject A kit characterized by containing the following.