Method for detecting gene mutation and method for differentiating somatic cell mutation from germ cell line mutation

JPWO2023027083A5Pending Publication Date: 2026-03-24
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Filing Date
2022-08-23
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Current methods for detecting genetic mutations in cancer tissues, particularly in low-tumor-content samples like diffuse gastric cancer, face challenges in accurately distinguishing between somatic and germline mutations due to limitations in tumor cell enrichment and the reliance on blood samples for germline mutation detection.

Method used

A method involving dissociation, separation, and sequencing of nucleic acid molecules from formalin-fixed paraffin-embedded (FFPE) tissue sections, using magnetic beads with specific ligands to enrich tumor cells and distinguish between somatic and germline mutations without the need for a blood sample, by analyzing mutant allele frequencies in tumor and residual fractions.

Benefits of technology

This approach increases the detection of genetic mutations and mutant allele frequencies, enabling accurate differentiation between somatic and germline mutations in FFPE tissue sections, even with low tumor cell proportions, thereby improving cancer treatment options.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

Provided are a method for detecting a gene mutation using an FFPE tissue section containing tumor cells regardless of the percentage of tumor cells, the method being capable of increasing the number of detectable gene mutations and mutant allele frequency, and a method capable of differentiating, even in the absence of blood samples, a somatic cell mutation from a germ cell line mutation. A method for detecting a gene mutation according to the present invention comprises: a dissociation step for dissociating a single cell group from a formalin-fixed paraffin-embedded tissue section containing tumor cells; a separation step for obtaining a tumor fraction containing the tumor cells from the single cell group; a collection step for collecting a nucleic acid molecule from the tumor fraction; and a sequencing step for subjecting the nucleic acid molecule to sequencing.
Need to check novelty before this filing date? Find Prior Art

Description

Methods for detecting gene mutations and distinguishing between somatic and germline mutations

[0001] The present invention relates to methods for detecting genetic mutations and for distinguishing between somatic and germline mutations.

[0002] In cancer treatment, limited genetic testing, such as companion diagnostics, provides cancer patients and clinicians with important information for effective drug selection. Recent large-scale analyses using next-generation sequencing (NGS) have revealed the relationship between genetic mutations and various cancers (Non-Patent Documents 1 to 3). Based on these findings, sequencing multiple target gene panels using NGS provides an opportunity for additional drug selection in clinical practice.

[0003] Alexandrov, L. B. et al. , Nature, 2013, Vol. 500, p. 415-421 Consortium, I. T. P. -C. A. o. W. G. , Nature, 2020, Vol. 578, p. 82-93 Nagashima, T. et al. , Cancer Sci, 2020, Vol. 111, p. 687-699

[0004] Detection of somatic mutations using NGS is affected by the tumor burden in tissue samples. Targeted panel sequencing is typically performed using formalin-fixed, paraffin-embedded (FFPE) tissue sections. When tumor content in FFPE tissue sections is low, tumor cell enrichment can be achieved by macrodissection. However, for cancers such as diffuse gastric cancer and lobular breast cancer, macrodissection is often unsuitable due to the diffuse nature of tumor cells. In particular, the tumor cell content in diffuse gastric cancer is often estimated to be less than 30%. Therefore, when sequencing targeted gene panels for various cancer types, alternative tumor cell enrichment methods other than macrodissection are needed to accurately detect mutations.

[0005] There are two standard targeted sequencing methods for detecting somatic mutations: one uses blood as a reference and the other uses public databases. While database-based methods have the advantage of being able to analyze FFPE tissue sections without blood as a reference, they inevitably carry the risk of falsely detecting alterations resulting from germline mutations as somatic mutations. Specifically, the accuracy of somatic mutation detection relies on public databases due to racial differences in single nucleotide polymorphisms (SNPs). Therefore, false positives increase in racial groups with insufficient SNP information. On the other hand, methods using blood from the same patient from which tissue was obtained can reliably determine germline mutations by subtracting mutations detected using blood as a reference, thereby enabling targeted sequencing to extract only somatic mutations. However, most archival samples stored as FFPE tissue sections are not paired with a blood reference that would enable somatic mutation detection based on targeted sequencing.

[0006] The present invention has been made in consideration of the above-mentioned problems, and aims to provide a method for detecting gene mutations that can improve the number of detectable gene mutations and mutant allele frequency using FFPE tissue sections containing tumor cells, regardless of the proportion of tumor cells, and a method that can distinguish between somatic mutations and germline mutations even without a blood sample.

[0007] The present inventors have conducted extensive research to solve the above-mentioned problems. As a result, they have found that the above-mentioned problems can be solved by dissociating a single cell population from an FFPE tissue section containing tumor cells, obtaining a tumor fraction containing the tumor cells from the single cell population, and thereby concentrating the tumor cells, thereby completing the present invention. More specifically, the present invention provides the following.

[0008] (1) A method for detecting a genetic mutation, comprising: a dissociation step of dissociating a single cell group from a formalin-fixed, paraffin-embedded tissue section containing tumor cells; a separation step of obtaining a tumor fraction containing the tumor cells from the single cell group; a recovery step of recovering nucleic acid molecules from the tumor fraction; and a sequencing step of sequencing the nucleic acid molecules.

[0009] (2) The method for detecting a gene mutation according to (1), wherein the thickness of the formalin-fixed, paraffin-embedded tissue section is 10 μm or more and 50 μm or less.

[0010] (3) The method for detecting a gene mutation according to (1) or (2), wherein the nucleic acid molecule is DNA.

[0011] (4) The method for detecting a gene mutation according to any one of (1) to (3), wherein the sequencing is next-generation sequencing.

[0012] (5) The method for detecting a gene mutation according to any one of (1) to (4), wherein the separation step comprises binding the tumor cells to magnetic beads and separating the magnetic beads bound to the tumor cells from cells other than the tumor cells by magnetic action, and the magnetic beads have a ligand that specifically binds to a biomolecule that is specifically present in the tumor cells.

[0013] (6) The method for detecting a gene mutation according to (5), wherein the biomolecule is at least one selected from the group consisting of cytokeratin and gene products of the following genes, and the ligand is an antibody against the biomolecule: HJURP gene, KIF2C gene, ASPN gene, GINS1 gene, NUSAP1 gene, IQGAP3 gene, CDK1 gene, TPX2 gene, CDT1 gene, MMP11 gene, MEX3A gene, TUBB3 gene, BIRC5 gene, HIST2H3A gene, CENPF gene, CCNB2 gene, TROAP gene, CDCA5 gene, KIAA0101 gene, UBE2C gene, AURKB gene, CKAP2L gene, and CEP55 gene. gene, EXO1 gene, KIF20A gene, CCNA2 gene, HIST1H2AL gene, ANLN gene, CENPA gene, TTK gene, ORC6 gene, SHCBP1 gene, FOXM1 gene, MELK gene, SPC25 gene, TOP2A gene, BUB1B gene, MAD2L1 gene, MND1 gene, KIFC1 gene, NUF2 gene, GTSE1 gene, E2F1 gene, BUB1 gene, DLGAP5 gene, KIF14 gene

[0014] (7) The method for detecting a gene mutation according to (5) or (6), wherein the biomolecule is cytokeratin and the ligand is an anti-cytokeratin antibody.

[0015] (8) A method for distinguishing between somatic mutations and germline mutations, comprising the dissociation step, separation step, recovery step, and sequencing step of the method for detecting a genetic mutation according to any one of (1) to (7), and further comprising: a second recovery step of recovering nucleic acid molecules from a residual fraction remaining after obtaining the tumor fraction in the separation step; a second sequencing step of sequencing the nucleic acid molecules recovered in the second recovery step; and an estimation step of estimating whether or not the target mutation detected in the sequencing step is a germline mutation, based on at least one of the mutant allele frequency obtained in the sequencing step and the mutant allele frequency obtained in the second sequencing step.

[0016] According to the present invention, it is possible to provide a method for detecting gene mutations that can increase the number of detectable gene mutations and mutant allele frequency using FFPE tissue sections containing tumor cells, regardless of the proportion of tumor cells, and a method that can distinguish between somatic mutations and germline mutations even without a blood sample.

[0017] 3A and 3B are optical microscope photographs showing diffuse gastric cancer (D1 and D2) and intestinal gastric cancer (S1 and S2) used in the examples. FFPE tissue sections stained with hematoxylin and eosin were used. The scale bar represents 2.5 mm. In the inset, areas with high tumor cell density are indicated by black arrows. The scale bar represents 100 μm. 3B is a graph showing the tumor cell mass in the unseparated sample, tumor fraction, and residue fraction obtained in the examples. 3C is a graph showing the average read depth. 3D is a graph showing estimated tumor content. 3E is a graph showing the effect of tumor cell enrichment on the detection of somatic mutations. Figure 4(A) is a graph showing the number of nonsynonymous mutations. Figure 4(B) is a Venn diagram showing the distribution of nonsynonymous mutations among unseparated samples, tumor fractions, and residue fractions. Figure 4(C) is a graph showing VAF (left) and read depth (right). * indicates p<0.01 / 3 (Welch's t-test with Bonferroni correction). Figure 4(D) is a graph showing the frequency of somatic mutations detected in diffuse gastric cancer and intestinal-type gastric cancer. Figure 4(E) is a graph showing the change in VAF of somatic mutations in unseparated samples, tumor fractions, and residue fractions. Figure 5(A) is a graph showing the distribution of VAF (left) and read depth (right). * indicates p<0.01. Figure 5(B) is a graph showing the results of comparing the VAF ratios of mutations common to the unseparated sample, tumor fraction, and residue fraction ((c) in Figure 4B) between germline mutations and somatic mutations. * indicates p<0.01. Figure 5(C) is a graph showing the receiver operating characteristic (ROC) curve for estimating the distinction between somatic mutations and germline mutations. Figure 6 is a heat map obtained by clustering expression levels in tumor areas based on the 21 types of tumors and 46 genes used in the examples.FIG. 7 shows the expression frequency and intracellular localization of expression in tumor tissues and normal tissues for the 46 genes used in the examples.

[0018] <Method for detecting genetic mutations> The method for detecting genetic mutations according to the present invention comprises: a dissociation step of dissociating a single cell population from an FFPE tissue section containing tumor cells; a separation step of obtaining a tumor fraction containing the tumor cells from the single cell population; a recovery step of recovering nucleic acid molecules from the tumor fraction; and a sequencing step of subjecting the nucleic acid molecules to sequencing. The method for detecting genetic mutations according to the present invention can increase the number of detectable genetic mutations and mutant allele frequency using an FFPE tissue section containing tumor cells, regardless of the proportion of tumor cells.

[0019] [Dissociation Step] In the dissociation step, single cell populations are dissociated from the FFPE tissue section containing tumor cells. The dissociation method is not particularly limited, and known methods can be used.

[0020] The thickness of the FFPE tissue section is not particularly limited and may be, for example, 10 μm to 50 μm. From the viewpoint of resource conservation and consistency with conventional methods, the thickness is preferably 10 μm to 20 μm, and more preferably 10 μm.

[0021] The proportion of tumor cells in an FFPE tissue section is not particularly limited. In the method for detecting gene mutations according to the present invention, even if the proportion is as small as, for example, 30% or less, preferably 15-25%, the number of detectable gene mutations and the mutant allele frequency can be improved. The proportion is measured as the proportion of the area occupied by tumor cells in the FFPE tissue section relative to the area occupied by the FFPE tissue section in an optical microscope photograph of the FFPE tissue section. The FFPE tissue section may be stained with hematoxylin and eosin, for example.

[0022] [Separation step] In the separation step, a tumor fraction containing the tumor cells is obtained from the single cell group. In this step, the tumor cells may be separated from the single cell group and the separated tumor cells may be collected to obtain the tumor fraction containing the tumor cells, or cells other than the tumor cells may be separated from the single cell group and the remaining residue may be collected to obtain the tumor fraction containing the tumor cells.

[0023] The method for separating tumor cells is not particularly limited, and known methods can be used. Examples of such separation methods include a method using a biomolecule that is specifically present in tumor cells. Specific examples include a method in which tumor cells are bound to a ligand that specifically binds to the biomolecule via the biomolecule, and the ligand bound to the tumor cells is then recovered. The biomolecules may be used alone or in combination of two or more. The ligands may be used alone or in combination of two or more.

[0024] In one embodiment, the biomolecule may be at least one selected from the group consisting of cytokeratin and gene products of the following genes: The gene product may be, for example, a protein; and the ligand may be, for example, an antibody against the biomolecule: HJURP gene, KIF2C gene, ASPN gene, GINS1 gene, NUSAP1 gene, IQGAP3 gene, CDK1 gene, TPX2 gene, CDT1 gene, MMP11 gene, MEX3A gene, TUBB3 gene, BIRC5 gene, HIST2H3A gene, CENPF gene, CCNB2 gene, TROAP gene, CDCA5 gene, KIAA0101 gene, UBE2C gene, AURKB gene, CKAP2L gene, and CEP55. gene, EXO1 gene, KIF20A gene, CCNA2 gene, HIST1H2AL gene, ANLN gene, CENPA gene, TTK gene, ORC6 gene, SHCBP1 gene, FOXM1 gene, MELK gene, SPC25 gene, TOP2A gene, BUB1B gene, MAD2L1 gene, MND1 gene, KIFC1 gene, NUF2 gene, GTSE1 gene, E2F1 gene, BUB1 gene, DLGAP5 gene, KIF14 gene

[0025] In another embodiment, the biological molecule includes, for example, a protein that is specifically present in tumor cells, such as cytokeratin, EpCAM, etc. The ligand includes, for example, an antibody against the protein.

[0026] The method for separating cells other than tumor cells is not particularly limited, and known methods can be used. Examples of such separation methods include a method using a biomolecule that is specifically present in cells other than tumor cells. Specific examples include a method in which cells other than tumor cells are bound to a ligand that specifically binds to the biomolecule via the biomolecule, and the ligand bound to the cells other than tumor cells is recovered. Examples of such biomolecules include proteins such as vimentin and fibronectin. Examples of such ligands include antibodies against the proteins.

[0027] In both the method for separating tumor cells and the method for separating cells other than tumor cells, the method for recovering the ligand is not particularly limited, and examples include a method in which the ligand is bound to an affinity carrier that specifically binds to the ligand and then recovered, or, if the ligand is bound to magnetic beads, a method in which the magnetic beads are recovered by magnetic action.

[0028] From the viewpoint of operability, etc., preferably, the separation step includes binding the tumor cells to magnetic beads and separating the magnetic beads bound to the tumor cells from cells other than tumor cells by magnetic action, the magnetic beads having a ligand that specifically binds to a biomolecule that is specifically present in the tumor cells. The biomolecule and the ligand are not particularly limited, and preferably, the biomolecule is at least one selected from the group consisting of cytokeratin and the gene product of the above-mentioned gene, and the ligand is an antibody against the biomolecule. More preferably, the biomolecule is cytokeratin, and the ligand is an anti-cytokeratin antibody. Specifically, commercially available products such as Anti-Cytokeratin MicroBeads (Miltenyi Biotec) can be used as the magnetic beads.

[0029] [Recovery step] In the recovery step, nucleic acid molecules are recovered from the tumor fraction. The method for recovering nucleic acid molecules is not particularly limited, and known methods can be used. The nucleic acid molecules are not particularly limited, and examples thereof include DNA and RNA, with DNA being preferred from the standpoint of operability, etc.

[0030] [Sequencing step] In the sequencing step, the nucleic acid molecule is subjected to sequencing. The sequencing method is not particularly limited, and examples thereof include NGS. The NGS method is not particularly limited, and known methods can be used.

[0031] <Method for Distinguishing Somatic Mutations from Germline Mutations> The method for distinguishing somatic mutations from germline mutations of the present invention comprises the dissociation step, separation step, recovery step, and sequencing step of the method for detecting a genetic mutation of the present invention, and further comprises: a second recovery step of recovering nucleic acid molecules from a residue fraction remaining after obtaining the tumor fraction in the separation step; a second sequencing step of sequencing the nucleic acid molecules recovered in the second recovery step; and an estimation step of estimating whether or not a target mutation detected in the sequencing step is a germline mutation based on at least one of the mutant allele frequency obtained in the sequencing step and the mutant allele frequency obtained in the second sequencing step. This method makes it possible to distinguish somatic mutations from germline mutations without a blood sample.

[0032] [Second Recovery Step] In the second recovery step, nucleic acid molecules are recovered from the residue fraction remaining after the tumor fraction is obtained in the separation step. Details of the second recovery step are the same as those of the recovery step in the method for detecting a gene mutation according to the present invention.

[0033] [Second Sequencing Step] In the second sequencing step, the nucleic acid molecules recovered in the second recovery step are subjected to sequencing. Details of the second sequencing step are the same as those of the sequencing step in the method for detecting a gene mutation according to the present invention.

[0034] In the estimation step, for a target mutation detected in the sequencing step, it is estimated whether the target mutation is a germline mutation based on at least one of the mutant allele frequency obtained in the sequencing step and the mutant allele frequency obtained in the second sequencing step. Specifically, the estimation step can be performed, for example, as described in the following embodiments 1 to 3.

[0035] (Embodiment 1) In embodiment 1, the estimation step includes calculating a VAF ratio, which is the ratio of the mutant allele frequency obtained in the sequencing step to the mutant allele frequency obtained in the second sequencing step, for a target mutation detected in the sequencing step, and estimating that the target mutation is a germline mutation if the VAF ratio is lower than a threshold. Note that the VAF ratio corresponds to a value expressed as (mutant allele frequency in the tumor fraction) / (mutant allele frequency in the residue fraction).

[0036] The threshold in embodiment 1 can be determined, for example, by analyzing in advance the relationship between the VAF ratio and the type of mutation (distinguishing between somatic mutations and germline mutations) for each race. Specifically, the threshold can be determined, for example, as follows: First, FFPE tissue sections and peripheral blood are collected from the same patient. Gene mutations are detected in the FFPE tissue sections using the method for detecting gene mutations according to the present invention, and the mutant allele frequencies are obtained for each of the tumor fraction and the residual fraction. Meanwhile, whole exome sequencing is performed on the peripheral blood to determine whether the gene mutation is a somatic mutation or a germline mutation. Based on these results, curves used as evaluation indices in binary classification, such as receiver operating characteristic (ROC) curves and precision-recall (PR) curves, are created for the VAF ratio and the type of mutation, assuming that the gene mutation is a somatic mutation as true, and the threshold can be determined.

[0037] (Embodiment 2) In embodiment 2, the estimation step includes calculating a VAF difference, which is the absolute value of the difference between the mutant allele frequency obtained in the sequencing step and the mutant allele frequency obtained in the second sequencing step, for a target mutation detected in the sequencing step, and estimating that the target mutation is a germline mutation if the VAF difference is lower than a threshold. Note that the VAF difference corresponds to the difference expressed as |(mutant allele frequency in the tumor fraction)-(mutant allele frequency in the residue fraction)|. The threshold in embodiment 2 can be determined in the same manner as the threshold in embodiment 1, except that the VAF difference is used instead of the VAF ratio.

[0038] (Embodiment 3) In embodiment 3, the estimation step includes estimating that a target mutation detected in the sequencing step is a germline mutation when the mutant allele frequency obtained in the second sequencing step is higher than a threshold. Note that the mutant allele frequency obtained in the second sequencing step corresponds to the mutant allele frequency in the residue fraction. The threshold in embodiment 3 can be determined in the same manner as the threshold in embodiment 1, except that the mutant allele frequency obtained in the second sequencing step is used instead of the VAF ratio.

[0039] The present invention will be explained in more detail below by showing examples, but the scope of the present invention is not limited to these examples.

[0040] Experimental Methods [Clinical Samples] Two types of diffuse gastric cancer and two types of intestinal gastric cancer were extracted from a Japanese pan-cancer cohort (Project HOPE), which included 5,521 tumor specimens. These samples were clinically and pathologically diagnosed by a pathologist after surgery. Tumors were immediately extracted from the surgical specimens after lesion resection at Shizuoka Cancer Center Hospital and preserved as FFPE tissue. Peripheral blood samples were also collected as paired controls to exclude germline mutations. Details of the experimental protocol are as described in the literature (Nagashima, T. et al. Cancer Sci 111, 687-699 (2020); Hatakeyama, K. et al. Cancer Sci 110, 2620-2628 (2019); Nagashima, T. et al. Biomed Res 37, 359-366 (2016); Shimoda, Y. et al. Biomed Res 37, 367-379 (2016); Urakami, K. et al. Biomed Res 37, 51-62, (2016); Ohshima, K. et al. Sci Rep 7, 641 (2017)). Briefly, DNA was extracted from tissue and peripheral blood samples using the QIAamp DNA Blood Mini Kit (Qiagen, Venlo, The Netherlands). Purified DNA was quantified using a NanoDrop and Qubit 2.0 Fluorometer (Thermo Fisher Scientific, Waltham, MA).

[0041] Dissociation and Suspension of FFPE Tissue Samples. FFPE tissue blocks of gastric cancer were cut into 10, 20, and 50 μm thick sections. These sections were delipidated by incubation in xylene three times for 10 minutes, and then rehydrated by incubation in 100% (twice), 70%, 50%, and 30% ethanol dilutions for 30 seconds. The hydration process was completed by incubation in deionized water for 30 seconds. After heat antigen retrieval according to the manufacturer's protocol, the delipidated samples were suspended using a gentleMACS Octo Dissociator with Heaters (Miltenyi Biotec, Bergisch Gladbach, Germany).

[0042] Cell Separation and Staining: Cells were labeled and separated automatically using an autoMACS Pro Separator (Miltenyi Biotec) according to the manufacturer's protocol. Specifically, cell suspensions derived from FFPE tissue sections were separated using Anti-Cytokeratin MicroBeads (Miltenyi Biotec). Cells were stained with anti-cytokeratin-FITC (clone REA831, Miltenyi Biotec), anti-vimentin-APC (clone REA409, Miltenyi Biotec), and CD235a (Glycophorin A)-PE (clone REA175, Miltenyi Biotec) antibodies. Nuclei were stained with DAPI Staining Solution (Miltenyi Biotec).

[0043] DNA isolation: DNA was extracted from FFPE tissues and peripheral blood samples using the GeneRead DNA FFPE Kit and QIAamp DNA blood Mini Kit (Qiagen), respectively. The purified DNA was quantified using a NanoDrop and Qubit 2.0 Fluorometer (Thermo Fisher Scientific). To confirm DNA quality, DIN was measured using a TapeStation (Agilent Technologies, Santa Clara, CA).

[0044] Targeted sequencing of gene panels: To perform targeted sequencing of genes in DNA isolated from FFPE tissues, a library consisting of 225 genes (listed in Table 1) was constructed using a hybridization-based enrichment protocol (SureSelect Custom panel, Agilent). A total of 2.427 Mb of the human genome, including 0.723 Mb of exonic regions of RefSeq genes, was covered with 55,765 biotinylated RNA oligomers (120 bp long). The binary raw data obtained from the sequencer was converted to sequence reads using bcl2fast1q (ver. 2.20, Illumina) and mapped to the reference human genome (UCSC hg19). To reduce false-positive findings, variants meeting any of the following criteria were excluded: (1) quality score <20, (2) depth of coverage <100, (3) depth of coverage for the alternate allele <5, (4) VAF <0.5%, and (5) not meeting the variant caller filtering criteria (the FILTER field in the VCF record is not "PASS"). After variant annotation, common SNPs with an allele frequency of 1% or higher were excluded in any of the following databases: (1) 1000 Genomes Project (Global or East Asia), (2) ExAC, and (3) gnomAD.Furthermore, mutations that are likely to affect protein structure, i.e., missense mutations, splice acceptor mutations, splice donor mutations, splice region mutations, stop-gain variants, stop-lost variants, stop-retained variants, 5'-untranslated region premature start codon gain variants, exon-loss variants, disruptive in-frame deletions, disruptive in-frame insertions, frameshift mutations, in-frame deletions, in-frame insertions, or start codon mutations, were extracted. To ensure reproducibility of sequencing, mutations with a VAF of 3% or higher were defined as valid mutations. Tumor content was estimated using the All-FIT algorithm based on tumor-only sequencing data (Loh, J. W. et al. Bioinformatics 36, 2173-2180, (2020)).

[0045]

[0046] Whole-exome sequencing: To accurately distinguish germline mutations without database-based predictions, we used the method described in Nagashima, T. et al. Cancer Sci 111, 687-699 (2020). Briefly, an exome library was constructed using the Ion Torrent AmpliSeq RDY Exome Kit (Thermo Fisher Scientific). The exome library provided 292,903 amplicons covering 57.7 Mb in the human genome and contained 34.8 Mb of exon sequences from 18,835 genes registered in RefSeq. To avoid sequencer- and amplicon-derived errors, any somatic mutations were manually inspected using Integrative Genomics Viewer (IGV), and somatic mutation candidates containing multiple nucleotide variations (approximately 1000 sites) were verified by Sanger sequencing.

[0047] Statistical Analysis: Significant differences in read depth and VAF (including VAF ratio) were determined using Welch's t-test. Bonferroni correction was performed for multiple comparisons. A P value of <0.01 was considered significant.

[0048] [Extraction of genes that can be used for cell isolation] In the cell isolation described above, cytokeratin was used as a biomolecule that is specifically present in tumor cells, and an anti-cytokeratin antibody was used as a ligand that specifically binds to the biomolecule. To identify the biomolecules other than cytokeratin, genes expressed without being affected by tumor heterogeneity were extracted by gene expression analysis. It is desirable that the candidate genes are not expressed in normal tissues (non-tumor tissues).

[0049] The specific extraction method is as follows: To extract genes expressed across cancer types, 21 types of tumors for which the applicant has expression information for both tumor and non-tumor areas were selected from tumors classified by OncoTree (Kundra et al., JCO Clinical Cancer Informatics 2021).

[0050] From the gene probes on a DNA microarray (Agilent Technologies), 20,869 protein-encoding genes were selected. Genes encoding hypothetical proteins, genes encoding putative proteins, and lincRNA detection probes were excluded. Using this DNA microarray, the expression levels in tumor and non-tumor areas of the 21 tumors were detected, and genes with an average (expression level in tumor) / (expression level in non-tumor) ratio of 2 or greater in 95% or more of the tumors, i.e., in 20 of the 21 tumors or in all 21 tumors, were extracted from the 20,869 genes.

[0051] Experimental Results Tumor Cell Enrichment Using Tissue Suspensions A total of 12 FFPE samples from four gastric cancer patients were obtained from the tissue bank of the Department of Diagnostic Pathology at Shizuoka Cancer Center, Shizuoka Prefecture. These samples included 10-, 20-, and 50-μm-thick FFPE tissue sections from two diffuse gastric cancers (D1 and D2) and two intestinal gastric cancers (S1 and S2) collected between 2014 and 2019 (Figure 1). The tumor cellularity estimated by the pathologist, i.e., the percentage of tumor cells in the FFPE tissue sections, was low in the diffuse gastric cancers (D1: 20%, D2: 20%) and high in the intestinal gastric cancers (S1: 60%, S2: 50%). These diffuse gastric cancers were considered unsuitable for macrodissection to enrich tumor cells in the FFPE tissue sections.

[0052] To increase the proportion of tumor cells in FFPE tissue sections from which DNA could be extracted, tumor cell enrichment was performed using tissue suspensions. Compared to unseparated samples, cells considered to be tumor cells (cytokeratin + , vimentin - ) were enriched in the tumor fraction, but were reduced in the residual fraction in both diffuse and intestinal gastric cancers (Figure 2). Furthermore, no differences in enrichment were observed depending on the thickness of the FFPE tissue sections. These results demonstrate that tumor cells expressing cytokeratin on the cell surface can be enriched from FFPE tissue sections of gastric cancer with low tumor burden.

[0053] [Confirming the Quality of Sequencing Samples] We investigated whether the quality of DNA extracted from tissue suspension samples was suitable for NGS. DNA degradation, DNA integrity number (DIN), and DNA concentration were used as indicators to determine whether the DNA quality was suitable for NGS (Figures 3(A) and 3(B)). These samples were used for library construction and NGS. The read depths of the unseparated and separated fractions were comparable (Figure 3(C)). Based on NGS, increased tumor content was observed in most tumor fraction samples (Figure 3(D)). These results suggest that NGS was performed appropriately using tumor fractions obtained from tissue suspension samples. Furthermore, although 50 μm-thick sections are recommended for preparing tissue suspensions, the quality of NGS reads was not affected by the thickness of the FFPE tissue sections. Based on these findings, we concluded that NGS can be performed using tissue suspensions prepared using 10 μm-thick FFPE tissue sections. Subsequent experiments were carried out using 10 μm thick sections.

[0054] Effect of Tumor Cell Enrichment To investigate whether tumor cell enrichment using tissue suspensions affects the detection of somatic mutations, nonsynonymous mutations were identified by targeted sequencing of a gene panel (225 genes listed in Table 1 were targeted). The number of mutations detected in the tumor fraction was greater than or equal to that detected in the unseparated sample, but the number of mutations detected in the residual fraction was lower than that detected in the unseparated sample (Figure 4(A)). Furthermore, 19% (25 / 133) of the mutations detected in the tumor fraction were specific to the tumor fraction (Figure 4(B)). These specific mutations (a) had significantly lower variant allele frequencies (VAFs) than those in (b) and (c) (see Figure 4(B) for mutations (a), (b), (c), and (d)), but there was no difference in read depth (Figure 4(C)). These results suggest that tumor cell enrichment using tissue suspensions can help identify somatic mutations that are not detected by conventional methods. Interestingly, tumor fraction-specific mutations (a) accounted for more than 30% of the mutations found in diffuse gastric cancer, suggesting that enrichment of tumor cells according to the present invention contributes to the detection of more mutations in this cancer type with low tumor burden (Fig. 4(D)). For mutations shared between the tumor fraction and unseparated samples, enrichment of tumor cells increased the VAF (Fig. 4(E)).

[0055] [Estimation of Germline Mutations Based on Differences Between Tumor and Residual Fractions] Mutations detected by sequencing of the targeted gene panel excluded germline mutations present in multiple databases. Therefore, SNPs not registered in databases, including those associated with racial differences, were identified as somatic mutations. To accurately distinguish between germline and somatic mutations, whole exome sequencing (WES) was performed using peripheral blood from patients who provided tumor tissue. Twenty-four mutations (18%) were identified as germline mutations by targeted panel sequencing (Tables 2-1 to 2-3). The VAF of somatic mutations revealed by WES of peripheral blood was significantly reduced in the unseparated sample and the residual fraction, but there was no difference in read depth (Figure 5(A)). Furthermore, one mutation common to both the unseparated sample and the residual fraction was included in the germline mutations revealed by WES of peripheral blood (Figure 4(B) (d)). These results raise the possibility that the VAF of germline mutations revealed by WES of peripheral blood may be independent of tumor burden in FFPE tissue sections. Based on this hypothesis, we compared the VAF ratio of shared mutations ((c) in Figure 4(B)) between germline and somatic mutations revealed by WES of peripheral blood. This ratio significantly increased for true somatic mutations (Figure 5(B)). Furthermore, a receiver operating characteristic (ROC) curve was constructed to distinguish somatic from germline mutations using the VAF ratio. The area under the curve (AUC) was 0.967, with a threshold VAF ratio of 0.668 (Figure 5(C)). These results demonstrate that germline mutations can be estimated by the VAF ratio using tumor and residue fractions obtained from FFPE tissue sections.

[0056] Summary: The examples demonstrated an increase in the number of detectable gene mutations and VAF. Furthermore, mutation analysis of DNA isolated from tumor and residual fractions allowed for the estimation of germline mutations without the need for blood samples, i.e., without using blood as a reference. This approach to enrich tumor cells not only increases the success rate of targeted panel sequencing, but also improves the accuracy of somatic mutation detection in specimens stored without blood samples, such as FFPE tissue sections.

[0057] [Extraction of genes that can be used for cell isolation] The following 46 genes were extracted from the 20,869 genes described above: HJURP gene, KIF2C gene, ASPN gene, GINS1 gene, NUSAP1 gene, IQGAP3 gene, CDK1 gene, TPX2 gene, CDT1 gene, MMP11 gene, MEX3A gene, TUBB3 gene, BIRC5 gene, HIST2H3A gene, CENPF gene, CCNB2 gene, TROAP gene, CDCA5 gene, KIAA0101 gene, UBE2C gene, AURKB gene, CKAP2L gene, and CEP55. gene, EXO1 gene, KIF20A gene, CCNA2 gene, HIST1H2AL gene, ANLN gene, CENPA gene, TTK gene, ORC6 gene, SHCBP1 gene, FOXM1 gene, MELK gene, SPC25 gene, TOP2A gene, BUB1B gene, MAD2L1 gene, MND1 gene, KIFC1 gene, NUF2 gene, GTSE1 gene, E2F1 gene, BUB1 gene, DLGAP5 gene, KIF14 gene

[0058] The expression levels in tumor areas were clustered using 21 types of tumors and 46 genes as axes, and a heat map was created (Figure 6). In Figure 6, for 46 genes, from HJURP to KIF14, the average (expression level in tumor areas) / (expression level in non-tumor areas) ratio was 2 or greater in more than 95% of 21 types of tumors, from CCRCC to COAD. In Figure 6, the expression levels in tumor areas were compared among the 46 genes whose expression levels in tumor areas were on average more than twice their expression levels in non-tumor areas in more than 95% of the tumors. In Figure 6, the HJURP to UBE2C genes tended to be relatively highly expressed in tumor areas, while the AURKB to KIF14 genes tended to be relatively low expressed in tumor areas. In Figure 6, for tumors ranging from LNET to COAD, 46 genes tended to be relatively highly expressed in the tumor area, and for tumors ranging from CCRCC to LUAD, 46 genes tended to be relatively low expressed in the tumor area.

[0059] The results for keratin genes (KRT7, KRT8, KRT18, and KRT19) are also shown in Figure 6. The expression levels of the 46 genes tended to be lower in tumor tissues than those of the keratin genes, but some genes were more highly expressed in tumor tissues than the keratin genes, depending on the tumor.

[0060] Using public databases, Protein Atlas (a database showing protein production due to gene expression by immunohistochemistry) was used to plot the expression frequencies of 46 genes in tumor tissues and normal tissues, and UniProt (a database of information on the intracellular localization of gene expression) was used to plot the intracellular localization of the expression of 46 genes (Figure 7). Figure 7 also shows the results for the above-mentioned keratin genes.

[0061] Figure 7 reveals the following: Of the 46 genes, several were found to be immunostained to the same extent or greater than keratin in tumor tissue (corresponding to the tumor area). Of the 46 genes, only a few were immunostained more strongly than keratin in normal tissue (corresponding to the non-tumor area). Therefore, using the 46 genes may enable more accurate separation of tumor cells and normal cells than using keratin. Furthermore, in Protein Atlas, there were genes whose protein production could not be observed by immunostaining in normal tissue (this may be due to the performance of the antibody). Furthermore, there were also genes in normal tissue (corresponding to the non-tumor area) that were not immunostained. The expression of the 46 genes tended to be particularly localized in the nucleus. From the above, it can be seen that the gene products of all 46 genes can be used as biomolecules in the separation process.

Claims

1. A dissociation process in which single cell groups are separated from formalin-fixed paraffin-embedded tissue sections containing tumor cells, A separation step to obtain a tumor fraction containing the tumor cells from the aforementioned single cell population, A recovery step for recovering nucleic acid molecules from the tumor fraction, A sequencing step of subjecting the nucleic acid molecule to sequencing, A method for detecting gene mutations, including those mentioned above.

2. The method for detecting gene mutations according to claim 1, wherein the thickness of the formalin-fixed paraffin-embedded tissue section is 10 μm or more and 50 μm or less.

3. The method for detecting gene mutations according to claim 1, wherein the nucleic acid molecule is DNA.

4. The method for detecting gene mutations according to claim 1, wherein the sequencing is next-generation sequencing.

5. The separation step includes binding the tumor cells to magnetic beads and separating the magnetic beads to which the tumor cells are bound from cells other than tumor cells by magnetic action. The method for detecting gene mutations according to claim 1, wherein the magnetic beads have ligands that specifically bind to biomolecules specifically present in the tumor cells.

6. The method for detecting gene mutations according to claim 5, wherein the biomolecule is at least one selected from the group consisting of cytokeratin and the gene product of the gene described below, and the ligand is an antibody against the biomolecule. HJURP gene, KIF2C gene, ASPN gene, GINS1 gene, NUSAP1 gene, IQGAP3 gene, CDK1 gene, TPX2 gene, CDT1 gene, MMP11 gene, MEX3A gene, TUBB3 gene, BIRC5 gene, HIST2H3A gene, CENPF gene, CCNB2 gene, TROAP gene, CDCA5 gene, KIAA0101 gene, UBE2C gene, AURKB gene, CKAP2L gene, CEP55 Genes: EXO1 gene, KIF20A gene, CCNA2 gene, HIST1H2AL gene, ANLN gene, CENPA gene, TTK gene, ORC6 gene, SHCBP1 gene, FOXM1 gene, MELK gene, SPC25 gene, TOP2A gene, BUB1B gene, MAD2L1 gene, MND1 gene, KIFC1 gene, NUF2 gene, GTSE1 gene, E2F1 gene, BUB1 gene, DLGAP5 gene, KIF14 gene

7. The method for detecting gene mutations according to claim 5, wherein the biomolecule is cytokeratin and the ligand is an anti-cytokeratin antibody.

8. A method for distinguishing between somatic mutations and germline mutations, A method for detecting gene mutations according to any one of claims 1 to 7, comprising a dissociation step, a separation step, a recovery step, and a sequencing step, Furthermore, In the separation step, a second recovery step is performed to recover nucleic acid molecules from the residual fraction remaining after obtaining the tumor fraction, A second sequencing step in which the nucleic acid molecules recovered in the second recovery step are subjected to sequencing, An estimation step is performed to estimate whether the target mutation detected in the sequencing step is a germline mutation, based on at least one of the mutation allele frequencies obtained in the sequencing step and the mutation allele frequencies obtained in the second sequencing step. A method that includes this.