Method for determining susceptibility to oral cancer and its application

The method uses an information processing device to analyze SYBR ratios of specific genes from oral rinse solutions, addressing the low accuracy and time-consuming issues of existing oral cancer diagnosis methods, enabling rapid and accurate susceptibility determination.

JP2026046923APending Publication Date: 2026-03-13SUDX BIOTEC CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-03
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing methods for diagnosing oral cancer, such as surgical biopsy, cytology, and oral rinse fluid analysis, suffer from low accuracy and are cumbersome or time-consuming, particularly the MS-MLPA method described in Patent Document 1.

Method used

A method using an information processing device that analyzes the SYBR ratios of specific genes (MLH1, BRCA2, and CDH13) from oral rinse solutions, employing machine learning to determine the likelihood of developing oral cancer based on the ratios of these genes' methylation frequencies, allowing for rapid and accurate diagnosis.

Benefits of technology

Enables rapid and accurate determination of oral cancer susceptibility with high specificity and simplicity, facilitating early detection and treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026046923000001_ABST
    Figure 2026046923000001_ABST
Patent Text Reader

Abstract

This method aims to provide a highly accurate and easy-to-use way to determine a person's risk of developing oral cancer in a short amount of time. [Solution] A method for determining the likelihood of a target developing oral cancer according to one embodiment of the present invention includes an acquisition step of acquiring the target's SYBR ratio obtained from the target's mouthwash, and a determination step of determining whether the target's likelihood of developing oral cancer is high or low using a machine learning model 21.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for determining the resistance to oral cancer and its use.

Background Art

[0002] Oral cancer that occurs in the oral region has an increasing incidence rate worldwide in recent years. In order to detect and treat oral cancer at an early stage, a definitive diagnosis by surgical biopsy is required. However, since the patient feels pain due to surgical biopsy, it is difficult to perform it frequently.

[0003] As methods for assisting the diagnosis of oral cancer, there are various techniques such as cytology, fluorescence scanning, saliva collection, and oral brushing. In addition, Patent Document 1 describes a method of using oral rinse fluid as a specimen to obtain data for the diagnosis of oral premalignancy.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, the above methods for assisting the diagnosis of oral cancer have the problem of low accuracy. In addition, the technique described in Patent Document 1 has the problem that since the MS-MLPA method is used, the operation is complicated and time-consuming.

[0006] One aspect of the present invention is in view of the above problems, and an object thereof is to realize a method for determining the resistance to oral cancer with high accuracy and in a short time by a simple operation.

Means for Solving the Problems

[0007] Therefore, one aspect of the present invention includes the following configuration. <1> A method for determining the likelihood of developing oral cancer, which is performed by an information processing device, Extracted from the target oral rinse solution, Gene A: A gene that is methylated infrequently, regardless of whether or not oral cancer is present. Gene B: A gene whose methylation frequency is more than twice as high in patients with oral cancer compared to patients without oral cancer. Gene C: A gene that is frequently methylated, regardless of whether or not oral cancer is present. The acquisition step involves obtaining the first SYBR ratio, which is the ratio of the first SYBR value to the second SYBR value, and the second SYBR ratio, which is the ratio of the first SYBR value to the third SYBR value, from the respective SYBR values ​​of genes A to C, namely the first SYBR value, the second SYBR value, and the third SYBR value. (1) An explanatory variable, including the first SYBR ratio and the second SYBR ratio, calculated using the respective SYBR values ​​of genes A to C extracted from each mouthwash of a first specimen suffering from oral cancer, is associated with an explanatory variable indicating the likelihood of developing oral cancer; and (2) An explanatory variable, including the respective SYBR values ​​of genes A to C extracted from each mouthwash of a second specimen not suffering from oral cancer, is associated with an explanatory variable indicating the likelihood of developing oral cancer; and the machine learning model generated by machine learning using training data is used to determine the likelihood of developing oral cancer in the subject based on the first SYBR ratio and the second SYBR ratio. A method for determining the likelihood of developing oral cancer. <2> The first SYBR ratio and the second SYBR ratio are the first SYBR-P and the second SYBR-P, respectively. <1> Methods used. <3> In the acquisition step described above, the product of the target first SYBR-P and the target second SYBR-P is further acquired. In the determination step, the explanatory variables of the machine learning model further include the product of the first SYBR-P and the second SYBR-P for each of the first and second samples, and the product of the first SYBR-P and the second SYBR-P of the target is further used to determine the degree of the target's susceptibility to oral cancer. <2> Methods used. <4> The subject and the specimen are of the same biological species. <1> Methods used. <5> The aforementioned gene A is the MLH1 gene, the aforementioned gene B is the BRCA2 gene, and the aforementioned gene C is the CDH13 gene. <1> Methods used. <6> For each of the first specimen, which was affected by oral cancer, and the second specimen, which was not affected by oral cancer, the following was extracted from the mouthwash: Gene A: A gene that is methylated infrequently, regardless of whether or not oral cancer is present. Gene B: A gene whose methylation frequency is more than twice as high in patients with oral cancer compared to patients without oral cancer. Gene C: A gene that is frequently methylated, regardless of whether or not oral cancer is present. The learning data acquisition step involves obtaining the first SYBR ratio, which is the ratio of the first SYBR value to the second SYBR value, and the second SYBR ratio, which is the ratio of the first SYBR value to the third SYBR value, from the SYBR values ​​of genes A to C, namely the first SYBR value, the second SYBR value, and the third SYBR value. The process includes a generation step of generating a machine learning model by performing machine learning using training data in which: (1) explanatory variables, including the first SYBR ratio and the second SYBR ratio, calculated using the respective SYBR values ​​of genes A to C extracted from each mouthwash of a first specimen affected by oral cancer, are associated with an objective variable indicating a low susceptibility to oral cancer; and (2) explanatory variables, including the first SYBR ratio and the second SYBR ratio, calculated using the respective SYBR values ​​of genes A to C extracted from each mouthwash of a second specimen not affected by oral cancer, are associated with an objective variable indicating a high susceptibility to oral cancer. How to create a machine learning model. <7> A method for determining the likelihood of developing oral cancer, which is performed by an information processing device, Extracted from the target oral rinse solution Gene A: A gene that is methylated infrequently, regardless of whether or not oral cancer is present. Gene B: A gene whose methylation frequency is more than twice as high in patients with oral cancer compared to patients without oral cancer. Gene C: A gene that is frequently methylated, regardless of whether or not oral cancer is present. Regarding this, the acquisition unit acquires the first SYBR ratio, which is the ratio of the first SYBR value to the second SYBR value, and the second SYBR ratio, which is the ratio of the first SYBR value to the third SYBR value, from the SYBR values ​​of genes A to C, namely the first SYBR value, the second SYBR value, and the third SYBR value, respectively. (1) An explanatory variable, including the first SYBR ratio and the second SYBR ratio, calculated using the respective SYBR values ​​of genes A to C extracted from the mouthwash of each first sample affected by oral cancer, is associated with an objective variable indicating a low likelihood of developing oral cancer; and (2) An explanatory variable, including the first SYBR ratio and the second SYBR ratio, calculated using the respective SYBR values ​​of genes A to C extracted from the mouthwash of each second sample affected by oral cancer, is associated with an objective variable indicating a high likelihood of developing oral cancer; and the machine learning model generated by machine learning using training data determines the likelihood of a subject developing oral cancer based on the first SYBR ratio and the second SYBR ratio. An information processing device for determining the likelihood of developing oral cancer. <8> <7> A control program for causing a computer to function as an information processing device as described above, wherein the control program causes the computer to function as the acquisition unit and the determination unit. <9> <8> A computer-readable recording medium containing the control program described above. [Effects of the Invention]

[0008] According to one aspect of the present invention, a method for determining the likelihood of developing oral cancer can be realized in a short time with high accuracy and simple operation. [Brief explanation of the drawing]

[0009] [Figure 1] This is a functional block diagram showing an example of the configuration of a determination device relating to one embodiment of the present invention. [Figure 2] This flowchart shows an example of a process flow for determining the likelihood of developing oral cancer using an information processing method according to one embodiment of the present invention. [Figure 3] This flowchart shows an example of the process for building a machine learning model related to one embodiment of the present invention. [Figure 4]It is a diagram showing an example of a ROC curve output by a machine learning model according to an embodiment of the present invention.

Embodiment for Implementing the Invention

[0010] 〔1. Information Processing Apparatus for Determining Resistance to Oral Cancer〕 Hereinafter, an information processing apparatus 10 according to an embodiment of the present invention will be described. The information processing apparatus 10 is an apparatus that determines the level of resistance to oral cancer of a target from the SYBR ratios of the target (the first SYBR ratio of the target and the second SYBR ratio of the target). Since the information processing apparatus 10 uses a machine learning model created based on data obtained from a first specimen actually suffering from oral cancer and a second specimen not suffering from oral cancer, it can accurately discriminate the resistance to oral cancer. In addition, since the information processing apparatus 10 uses the values obtained by real-time PCR, it can perform determination in a short time with a simple operation. [[ID=^{12}]]

[0011] In this specification, the "SYBR ratio" means Gene A: A gene with a low frequency of methylation regardless of the presence or absence of oral cancer, Gene B: A gene with a methylation frequency in patients suffering from oral cancer that is at least twice as high as the methylation frequency in patients not suffering from oral cancer, Gene C: A gene with a high frequency of methylation regardless of the presence or absence of oral cancer a value calculated based on the SYBR values calculated for each of them. More specifically, the ratio of the SYBR value of gene B to the SYBR value of gene A is defined as the first SYBR ratio, and the ratio of the SYBR value of gene C to the SYBR value of gene A is defined as the second SYBR ratio.

[0012] Also, in this specification, the first SYBR ratio and the second SYBR ratio calculated from the oral rinse fluid of the target are collectively referred to as the "SYBR ratio of the target". Also, in this specification, the first SYBR ratio and the second SYBR ratio calculated from the oral rinse fluid of each of the first specimen and the second specimen are collectively referred to as the "SYBR ratio of the specimen".

[0013] In this specification, the "SYBR value" is a value obtained by performing melting curve analysis based on the decrease in the fluorescence value of SYBR Green for the genes A to C described above, and integrating a specific range of the melting curve.

[0014] More specifically, for example, when using Thermal Cycler Dice® Type III (manufactured by Takara Bio Inc.), first, 45 cycles of PCR are performed for each gene, then the temperature is increased by 0.5°C every 30 seconds from 60°C to 95°C, and melting curve analysis is performed based on the decrease in the fluorescence value of SYBR Green. Here, the melting curve can be converted into a melting peak by plotting the negative first derivative of the fluorescence value (-dF / dT) as a function of temperature. Next, based on the melting peak, the melting temperature (Tm value) at which the absolute value of -dF / dT is maximized is determined for each gene. Furthermore, the Tm value ±0.5°C and the Tm value +1.0°C or -1.0°C (whichever is greater in absolute value of -dF / dT) are determined. The SYBR value is calculated by summing the absolute values ​​of -dF / dT at these four temperature points (Tm, Tm+0.5℃, Tm-0.5℃, Tm+1.0℃ or -1.0℃) and then integrating the sum.

[0015] Furthermore, when using the Lightcycler® 96 (manufactured by Roche), first, 45 cycles of PCR are performed for each gene, and then the temperature is raised by 0.3°C every 30 seconds from 60°C to 95°C, and melting curve analysis is performed. After converting the melting curve to a melting peak, the Tm value is determined for each gene, and the Tm value ±0.5°C and Tm value +1.0°C or -1.0°C (whichever is larger in absolute value of -dF / dT) are determined. The absolute values ​​of -dF / dT at these four temperature points (Tm, Tm+0.5°C, Tm-0.5°C, Tm+1.0°C or -1.0°C) are summed, and the integral value calculated is the SYBR value. The SYBR value can be calculated regardless of the type of instrument.

[0016] In this specification, "low frequency of methylation" means that after examining a specific gene in multiple individuals (preferably 60 or more), determining the total number of methylated bases in the gene, and then dividing this total by the number of individuals, the result is 0.2 or less. "High frequency of methylation" means that after examining a specific gene in multiple individuals (preferably 60 or more), determining the total number of methylated bases in the gene, and then dividing this total by the number of individuals, the result is 1.0 or more. Low frequency and high frequency of methylation can be confirmed, for example, by the MS-MLPA method.

[0017] The statement "The frequency of methylation in patients with oral cancer is more than twice as high as the frequency of methylation in patients without oral cancer" means that the value obtained for patients with oral cancer is more than twice as high as the value obtained for patients without oral cancer.

[0018] For example, gene A can be MLH1, RASSF1, or PTEN. For example, gene B can be BRCA2 or DAPK1. For example, gene C can be CDH13, CDKN2B, or CD44. Among these, gene A is preferably the MLH1 gene, gene B is the BRCA2 gene, and gene C is the CDH13 gene.

[0019] Gene B is a gene in which methylation is more than twice as frequent in patients with oral cancer compared to patients without oral cancer. Therefore, a machine learning model created using the SYBR ratio calculated based on the SYBR values ​​of genes A to C mentioned above can accurately detect subjects with high frequency of gene B methylation (i.e., subjects with a low risk of developing oral cancer). It is remarkable that this method can detect subjects with a low risk of developing oral cancer more simply and accurately than conventional cytology.

[0020] In one embodiment, the target SYBR ratio is the target SYBR-P, and the sample SYBR ratio is the sample SYBR-P. SYBR-P is the p-value obtained by comparing the first SYBR ratio and the second SYBR ratio with the first SYBR ratio and the second SYBR ratio of the positive control described above for each gene, and determining the significant difference by a t-test. That is, for example, the target SYBR-P includes the p-value (first SYBR-P) obtained when determining the significant difference between the SYBR ratio of target gene 2 and the SYBR ratio of positive control gene 22, and the p-value (second SYBR-P) between the SYBR ratio of target gene 3 and the SYBR ratio of positive control gene 3. The sample SYBR ratio may also be, for example, a SYBR-P that includes a p-value, similar to the target SYBR ratio. The above configuration improves the accuracy of the determination.

[0021] In one embodiment, the SYBR ratio of the target and the SYBR ratio of the sample may each further include the product of the SYBR-P of gene 2 (first SYBR-P) and the SYBR-P of gene 3 (second SYBR-P). This configuration further improves the accuracy of the determination.

[0022] Figure 1 is a functional block diagram showing an example of the configuration of the information processing device 10. The information processing device 10 includes a control unit 1 that controls all parts of the information processing device 10, a storage unit 2 that stores various data used by the information processing device 10, an input unit 11, and an output unit 15, but is not limited to this configuration. For example, the storage unit 2 and the input unit 11 may be external devices attached to the information processing device 10. In that case, the input unit 11 may be an input device such as a laptop computer or tablet connected to the information processing device 10 by wire or wireless connection. The output unit 15 and the output control unit 14 may be a communication unit and a communication control unit that transmit the output results to another device or network.

[0023] The input unit 11 is for receiving various input operations from the user and may be, for example, a keyboard, mouse, touch panel, etc. The input unit 11 may be used to input information that includes at least the SYBR ratio of the subject. The information input via the input unit 11 may further include the type of oral cancer and the species of the subject or specimen.

[0024] The output unit 15 outputs the degree of susceptibility to oral cancer. The output mode of the output unit 15 is not particularly limited and may be, for example, a display output, a printed output, or an audio output.

[0025] <Control Unit 1> The control unit 1 comprises an acquisition unit 12, a determination unit 13, and an output control unit 14. Furthermore, some of the blocks included in the control unit 1 may be delegated to other devices capable of communicating with the information processing device 10, and those blocks may be omitted from the control unit 1.

[0026] The acquisition unit 12 acquires data input from the input unit 11. The acquisition unit 12 acquires at least the SYBR ratio of the target. The acquisition unit 12 may further acquire data on the type of oral cancer and the species of the target or specimen. The acquisition unit 12 may store the acquired data in the storage unit 2. The acquisition unit 12 may perform preprocessing on the acquired data.

[0027] The determination unit 13 inputs the target's SYBR ratio 22 into a trained machine learning model 21 to determine the likelihood of the target developing oral cancer. The determination unit 13 may store the determination result 23 in the storage unit 2. In this specification, "high likelihood of developing" means that the possibility of developing or having developed oral cancer is low. Also, "low likelihood of developing" means that the possibility of developing or having developed oral cancer is high. In one embodiment, the determination unit 13 may determine that the determination result 23 for oral cancer is one of two options, "prone to developing" or "not prone to developing," or one of three or more options, such as "very prone to developing," "moderate," or "very unprone to developing."

[0028] In one embodiment, the determination unit 13 may determine whether or not there is a risk of oral cancer in the target.

[0029] In one embodiment, the determination unit 13 may determine the likelihood of developing oral cancer based on a value output by a mathematical formula calculated by the machine learning model 21 based on the sample's SYBR ratio and data related to the sample's likelihood of developing oral cancer.

[0030] The output control unit 14 causes the output unit 15 to output the estimation result output by the determination unit 13.

[0031] <Storage section 2> The memory unit 2 may store the machine learning model 21. Additionally, it may store the target SYBR ratio 22 and the judgment result 23 as needed.

[0032] The machine learning model 21 learns the relationship between the SYBR ratio of a sample, obtained from the sample's oral rinse solution, and the sample's susceptibility to oral cancer. This allows the machine learning model 21 to determine the judgment result 23 from the target SYBR ratio 22.

[0033] Conventional diagnostic methods have involved painful surgical biopsies for patients, low-accuracy cytology, and low-specificity fluorescence imaging. Furthermore, highly accurate methods are cumbersome and time-consuming to obtain results. In contrast, the information processing device 10 can provide diagnostic results quickly, using a simple method with high accuracy and specificity. By providing diagnostic results quickly and simply, the information processing device 10 enables early detection and treatment of rapidly progressing oral cancer.

[0034] [2. Method for determining the likelihood of developing oral cancer] A method for determining susceptibility to oral cancer according to one embodiment of the present invention (hereinafter also simply referred to as this determination method) involves obtaining the first SYBR ratio, which is the ratio of the first SYBR value to the second SYBR value, and the second SYBR ratio, which is the ratio of the first SYBR value to the third SYBR value, from the SYBR values ​​of genes A to C, respectively: (1) Oral The method includes: (1) a machine learning model generated by machine learning using training data, in which an explanatory variable including the first SYBR ratio and the second SYBR ratio, calculated using the respective SYBR values ​​of genes A to C extracted from each mouth rinse fluid of a first specimen suffering from oral cancer, is associated with an objective variable indicating the likelihood of developing oral cancer; and (2) an explanatory variable including the first SYBR ratio and the second SYBR ratio, calculated using the respective SYBR values ​​of genes A to C extracted from each mouth rinse fluid of a second specimen not suffering from oral cancer, is associated with an objective variable indicating the likelihood of developing oral cancer. The method includes a determination step of determining the likelihood of developing oral cancer in a subject based on the first SYBR ratio and the second SYBR ratio of the subject. This determination method is a method executed by an information processing device. In other words, this determination method can also be said to be an information processing method.

[0035] This determination method will be explained based on Figure 2. Figure 2 is a flowchart outlining this determination method. In the following explanation, MLH1 is selected as gene A, BRCA2 as gene B, and CDH13 as gene C.

[0036] In step S1, the acquisition unit 12 acquires the target SYBR ratio 22 (acquisition step). The target SYBR ratio 22 is obtained from the SYBR values ​​of the MLH1 gene, CDH13 gene, and BRCA2 gene extracted from the target oral rinse solution.

[0037] In this specification, "oral rinse solution" means a liquid sample obtained by a subject gargling with an appropriate amount of water or aqueous solution in the mouth. The water or aqueous solution is not particularly limited, but may be, for example, physiological saline, distilled water, or general tap water. The oral rinse solution is particularly suitable for non-invasive diagnosis because it contains exfoliated cells from the entire oral mucosa, is easy to collect, and is easy to process for detection.

[0038] This determination method may include a step S2 in which the acquisition unit 12 performs preprocessing on the target SYBR ratio 22. The preprocessing is not particularly limited, but examples include averaging, standardization, and outlier removal. By performing preprocessing on the target SYBR ratio 22, the accuracy of the determination result 23 can be improved.

[0039] In step S3, the determination unit 13 inputs the target SYBR ratio 22 into the machine learning model 21 to determine the level of the target's susceptibility to oral cancer. At this time, the machine learning model 21 is obtained by machine learning using the SYBR ratio of the sample and the data on the level of the sample's susceptibility to oral cancer, which are obtained from the oral rinse solution of the sample.

[0040] The subjects and specimens described above are not particularly limited as long as they are species that can develop oral cancer, but may include mammals, for example. Preferably, the subjects and specimens are of the same species. In one embodiment, the species of subjects and specimens are humans. The subjects are preferably patients who may have abnormalities in their oral cavity, and patients undergoing health checkups. Preferably, the specimens include both specimens with oral cancer (first specimen) and specimens without oral cancer (second specimen). This improves the accuracy of the diagnosis.

[0041] This determination method may include step S4, in which the output control unit 14 outputs the determination result 23. In step S4, the information processing device causes the output device to output the determination result 23 as data relating to the likelihood of developing oral cancer in the subject. By having an output step, this determination method allows for easy confirmation of the determination result 23. In step S4, the SYBR ratio 22 of the subject may be further output.

[0042] The inventors investigated various literature, research data, the genes of oral cancer patients, and the genes of individuals without oral cancer, and selected eight genes with high and low frequencies of DNA methylation, or those in which DNA is not methylated. From the base sequences of these eight genes, primer sets were created that include regions where DNA methylation can occur. In addition, two primer sets were created for MLH1. The created primer sets are as follows. RASSF1 (Forward: SEQ ID NO: 1, Reverse: SEQ ID NO: 2), CD44 (Forward: SEQ ID NO: 3, Reverse: SEQ ID NO: 4), BRCA2 (Forward: SEQ ID NO: 5, Reverse: SEQ ID NO: 6), CDKN1B (Forward: SEQ ID NO: 7, Reverse: SEQ ID NO: 8), CDH13 (Forward: SEQ ID NO: 9, Reverse: SEQ ID NO: 10), CDKN2B (Forward: SEQ ID NO: 11, Reverse: SEQ ID NO: 12), PTEN (Forward: SEQ ID NO: 13, Reverse: SEQ ID NO: 14), MLH1-1 (Forward: SEQ ID NO: 15, Reverse: SEQ ID NO: 16), MLH1-2 (Forward: SEQ ID NO: 17, Reverse: SEQ ID NO: 18).

[0043] To obtain a positive control cDNA, individuals in whom methylation was not occurring were identified for the eight genes mentioned above, as well as 16 additional genes (RARB, ESR1, DAPK1, KLLN, ATM, VHL, TP73, CASP8, FHIT, APC, CDKN2A, GSTP1, CADM1, HIC1, BRCA1, TIMP3). Oral rinse fluid was collected from these individuals, and PCR was performed using the primer set described above with the total DNA extracted from the oral rinse fluid as a template. A plasmid was created by incorporating the resulting cDNA into a vector, and this plasmid was used as a positive control for quantitative PCR.

[0044] In quantitative real-time PCR using the intercalator method, when a specific gene is amplified, the solution in which the PCR reaction is progressing gradually becomes fluorescent due to fluorescent molecules intercalated within the DNA double helix. Generally, the amplification of a specific gene is indicated by the intensity of this fluorescence. More specifically, a specific fluorescence intensity is used as a threshold, the number of PCR cycles exceeding this threshold is defined as the Ct value, and the concentration of the original gene is estimated based on this Ct value. However, subtle changes in gene concentration are difficult to discern from the Ct value alone. Therefore, the inventors have diligently investigated and analyzed gene concentration using the melting temperature curve of the PCR product, rather than the Ct value.

[0045] Gene concentration analysis was performed using a Thermal Cycler Dice® Type III (manufactured by Takara Bio Inc.). After 45 cycles of PCR, the temperature was increased by 0.5°C every 30 seconds from 60°C to 95°C, and melting curve analysis was performed based on the decrease in fluorescence value of SYBR Green. The melting curve is converted to a melting peak by plotting the negative first derivative (-dF / dT) of the fluorescence value as a function of temperature. First, for each gene, the melting temperature (Tm value) at which the absolute value of -dF / dT is maximized was determined. Furthermore, Tm value ±0.5°C and Tm value +1.0°C or -1.0°C (whichever has the larger absolute value of -dF / dT) were determined, and the absolute values ​​of -dF / dT at these four temperature points (Tm, Tm+0.5°C, Tm-0.5°C, Tm+1.0°C or -1.0°C) were summed and then integrated to calculate the integral value (SYBR value).

[0046] As described above, quantitative real-time PCR was performed using the intercalator method (SYBR-G method) with the plasmid and primer set mentioned above. Genes with an error of less than 12% in the SYBR value in the PCR performed simultaneously were selected. Specifically, MLH1 was selected as gene A, BRCA2 as gene B, and CDH13 as gene C from the genes for which the primer set was created.

[0047] In the following, the ratio of the SYBR value of the BRCA2 gene to the SYBR value of the MLH1 gene was defined as the BRCA2 SYBR ratio (first SYBR ratio), and the ratio of the SYBR value of the CDH13 gene to the SYBR value of the MLH1 gene was defined as the CDH13 SYBR ratio (second SYBR ratio).

[0048] In addition to the above SYBR ratio, plasmids containing only the MLH1 gene sequence, plasmids containing only the BRCA2 gene sequence, and plasmids containing only the CDH13 gene sequence may be subjected to quantitative real-time PCR as positive controls, with their copy numbers kept constant, and the SYBR ratio of the positive controls may be determined in the same manner. In one embodiment, the SYBR ratio of the sample and the SYBR ratio of the target may include the SYBR ratio of the positive control described above.

[0049] In one embodiment, the target SYBR-P includes the p-value (first SYBR-P) obtained when the significant difference between the SYBR ratio of BRCA2 and the SYBR ratio of the positive control BRCA2 is determined, and the p-value (second SYBR-P) obtained between the SYBR ratio of CDH13 and the SYBR ratio of the positive control CDH13. The sample SYBR ratio may also include a p-value, similar to the target SYBR ratio. This configuration improves the accuracy of the determination.

[0050] Oral cancer is not limited to any cancer that occurs in the oral cavity, but includes oral mucosal lesions such as oral occult malignancies (OPMDs), including precancerous lesions of oral cancer such as erythroplakia and leukoplakia, squamous cell carcinoma (SCC), and carcinoma in situ (CIS).

[0051] The subjects and specimens described above are not particularly limited as long as they are species that can develop oral cancer, but may include mammals, for example. Preferably, the subjects and specimens are of the same species. In one embodiment, the species of subjects and specimens are humans. The subjects are preferably patients who may have abnormalities in their oral cavity and patients undergoing health checkups. The specimens preferably include both patients diagnosed with oral cancer and healthy humans. This improves the accuracy of the diagnosis.

[0052] [3. How to create a machine learning model] A method for creating a machine learning model 21 according to one embodiment of the present invention includes a training data acquisition step of acquiring the SYBR ratio of a sample and data on the sample's susceptibility to oral cancer, obtained from the oral rinse fluid of the sample, and a generation step of generating a machine learning model by performing machine learning using the SYBR ratio of the sample as an explanatory variable and the data on the sample's susceptibility to oral cancer as an objective variable, wherein the SYBR ratio of the sample is the ratio of the SYBR value of the BRCA2 gene to the SYBR value of the MLH1 gene, and the ratio of the SYBR value of the CDH13 gene to the SYBR value of the MLH1 gene (wherein the MLH1 gene, BRCA gene, and CDH13 gene are genes extracted from the oral rinse fluid of the sample).

[0053] The method for creating the machine learning model 21 will be explained based on Figure 3. Figure 3 is a flowchart showing an example of how to create a machine learning model.

[0054] Step S11 is a training data acquisition step in which the SYBR ratio and data regarding the presence or absence of oral cancer are obtained for the first and second samples, respectively.

[0055] Step S12 involves performing machine learning using the SYBR ratios of the first and second samples as explanatory variables, and data that correlates the presence or absence of oral cancer with the likelihood of developing oral cancer as the target variable. Step S12 may be performed by a machine learning device equipped with a learning unit for performing the machine learning described above.

[0056] In step S12, known machine learning algorithms can be used as machine learning methods. Examples of machine learning algorithms that can be used to generate the trained model 21 include k-nearest neighbor method, logistic regression, support vector machines, random forests, and neural networks.

[0057] The data on the likelihood of developing oral cancer in the specimen used in step S12 is not particularly limited, as long as it is data that can be used to determine the probability of developing oral cancer in the specimen. For example, it may be data related to the risk of developing oral cancer based on genetic testing, or data related to the presence or absence of oral cancer in the specimen diagnosed by surgical biopsy, etc.

[0058] More specifically, the following methods 1 to 3 can be used to perform machine learning in step S12. Among these, method 3 is preferable because it improves the accuracy of the judgment. In methods 1 to 3 below, n is any even natural number greater than or equal to 6. Method 1: Compare the SYBR ratio between cancer patients and healthy individuals. In this case, real-time PCR is performed n times on the same sample, and the average of (n-2) data points, excluding the maximum and minimum values ​​of the SYBR ratio, is used as the SYBR ratio. If the standard error of the (n-2) data points is not less than 12%, the real-time PCR is repeated (n / 2) times, and the average of the five data points obtained by excluding the two maximum and two minimum values ​​from a total of 1.5 × n data points is used. If the standard error is still not less than 12%, the measurement is repeated. Method 2: The SYBR ratio is calculated in the same way as in Method 1. Based on the SYBR ratio, the SYBR-P values ​​for BRCA2 and CDH13 are used as the sample SYBR ratios to create a machine learning model. Method 3: In addition to Method 2, the product of SYBR-P is used as the SYBR ratio.

[0059] In step S12, a threshold (cutoff value) may be determined. The cutoff value may be automatically determined when the machine learning model 21 is created, or it may be determined by the creator of the machine learning model 21 based on the ROC (Receiver Operating Characteristic) curve output when the machine learning model 21 is created. Figure 4 shows an ROC curve according to one embodiment of the present invention. The cutoff value may be determined by selecting any point on the ROC curve that yields the desired sensitivity and 1-specificity values.

[0060] If the value calculated by the machine learning model 21 exceeds the cutoff value, the determination unit 13 determines that the likelihood of developing the target oral cancer is low, and if it falls below the cutoff value, the determination unit 13 determines that the likelihood of developing the target oral cancer is high. In one embodiment, if the value calculated by the machine learning model 21 is the same as the cutoff value, the creator of the machine learning model may decide whether the determination unit 13 determines it as high or low.

[0061] [Examples of implementation using software] The functions of the information processing device 10 (hereinafter referred to as "the device") are programs that cause the device to function as a computer, and these programs can be realized by programs that cause each control block of the device (especially each part included in the control unit 1) to function as a computer.

[0062] In this case, the device includes a computer having at least one control device (e.g., a processor) and at least one storage device (e.g., memory) as hardware for executing the program. By executing the program using this control device and storage device, the functions described in each of the embodiments are realized.

[0063] The above program may be recorded on one or more computer-readable recording media, not temporary ones. These recording media may or may not be provided by the above device. In the latter case, the program may be supplied to the above device via any wired or wireless transmission medium.

[0064] Furthermore, some or all of the functions of each of the above control blocks can also be realized by logic circuits. For example, an integrated circuit in which logic circuits functioning as each of the above control blocks are formed is also included in the scope of the present invention. In addition, it is also possible to realize the functions of each of the above control blocks by, for example, a quantum computer.

[0065] Furthermore, each process described in the above embodiments may be performed by AI (Artificial Intelligence). In this case, the AI ​​may operate on the control device described above, or it may operate on other devices (for example, an edge computer or a cloud server).

[0066] The present invention is not limited to the embodiments described above, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention. [Examples]

[0067] One embodiment of the present invention is described below.

[0068] [Test Method] Between 2022 and 2023, 60 patients who visited the Department of Oral Surgery at Kagoshima University Hospital were classified into two groups: a low-risk group (30 patients with a confirmed diagnosis of squamous cell carcinoma (SCC) or carcinoma in situ (CIS) based on pathological diagnosis) and a high-risk group (30 healthy individuals without oral mucosal disease). Two independent cohorts were created: a training cohort and a test cohort. Sample size was determined using G*Power (version 3.1). The training cohort consisted of data from 20 patients and 20 healthy individuals. The test cohort consisted of data from 10 patients and 10 healthy individuals. The sample size for the test cohort was calculated based on the area under the receiver operating characteristic (ROC) curve (AUC) obtained from the training results of the training cohort. In the training cohort, 15 patients had SCC and 5 had CIS. In the test cohort, 7 patients had SCC and 3 had CIS. There were no statistically significant differences in age, sex, or smoking status between patients and healthy controls in each cohort. The sample was the total solution of 20 mL of sterile purified water in which the subject rinsed their mouth for 30 seconds.

[0069] Oral rinse solution samples (20 mL) obtained from subjects were centrifuged at 25°C and 500 G for 5 minutes, and the cell pellet in the oral rinse solution was separated as a precipitate. 200 μL of phosphate-buffered saline was added to the cell pellet to suspend it, and DNA was extracted and purified using the Dneasy Blood and Tissue Kit (QIAGEN) according to the attached protocol. The extracted and purified DNA was quantified using NanoDrop (ThermoFisher).

[0070] Quantitative real-time PCR was performed using purified DNA obtained by the intercalator method (SYBR Green method). A Thermal Cycler Dice III Real Time System (Takara Bio Inc.) was used as the PCR analyzer. The PCR protocol involved an initial activation step of 20 seconds at 95°C, followed by 45 cycles of denaturation and extension at 95°C for 3 seconds and 60°C for 30 seconds. Subsequently, melting curve analysis of the PCR product was performed (denaturation at 95°C for 15 seconds, followed by a stepwise temperature increase of 0.5°C from 60°C to 95°C every 30 seconds).

[0071] Statistical analysis was performed using IBM SPSS Statistics version 26 (IBM Corporation). The area under the ROC curve (AUC) was calculated from the ROC curve, with the detection of malignant tumors (SCC, CIS) as the endpoint. Based on the sensitivity and specificity obtained from the ROC curve, a useful cutoff value for diagnosis was determined. Fisher's exact test was used to evaluate diagnostic performance, with a significance level of P < 0.05. Sample size was determined using G*Power (version 3.1).

[0072] [Indicator A] Using logistic regression analysis, a judgment formula (machine learning model A) was created to detect the low susceptibility group according to Method 1 described above. At this time, a judgment formula was created using the SYBR ratio of BRCA2 and the SYBR ratio of CDH13 as test variables, and this was designated as index A. Index A = 3.57 + (3.627 × SYBR ratio of BRCA2) + (-1.268 × SYBR ratio of CDH13) ROC analysis was performed using index A, and Fisher's exact test was further used to determine the cutoff value (0.454). The diagnostic performance of index A was as follows: AUC = 0.780, sensitivity = 85.0%, specificity = 70.0%, positive predictive value (PPV) = 73.9%, and negative predictive value (NPV) = 82.4%.

[0073] To evaluate the reproducibility of detecting the low-risk group using Indicator A, we validated Indicator A using a cohort consisting of test data (10 cases from the low-risk group and 10 cases from the high-risk group). The diagnostic performance was AUC=0.830, sensitivity=80.0%, specificity=60.0%, PPV=66.7%, and NPV=50%. Therefore, it was confirmed that Indicator A can reproducibly detect the low-risk group. Furthermore, Indicator A also showed high diagnostic performance in detecting the malignant group in the test data.

[0074] [Indicator B] Using logistic regression analysis, a judgment formula was created to detect the low-susceptibility group according to (Method 2) described above. In this case, a judgment formula was created using two items, SYBR-P of BRCA2 and SYBR-P of CDH13, as test variables, and this was designated as Index B (Machine Learning Model B). Index B = 0.6612 + (-0.9593 × SYBR-P of BRCA2) + (-0.344 × SYBR-P of CDH13) ROC analysis was performed using index B, and the cutoff value (0.514) was determined from Fisher's exact test. The diagnostic performance of index B was as follows: AUC = 0.808, sensitivity = 90.0%, specificity = 65.0%, positive predictive value (PPV) = 72.0%, and negative predictive value (NPV) = 86.7%.

[0075] To evaluate the reproducibility of detecting the low-risk group using Indicator B, we validated Indicator B using a cohort consisting of test data (10 cases from the low-risk group and 10 cases from the high-risk group). The diagnostic performance was AUC=0.820, sensitivity=90.0%, specificity=80.0%, PPV=81.8%, and NPV=88.9%. Therefore, it was confirmed that Indicator B can reproducibly detect the low-risk group. Furthermore, Indicator B also showed high diagnostic performance in detecting the malignant group in the test data.

[0076] [Indicator C] Using logistic regression analysis, a judgment formula was created to detect the low-susceptibility group according to Method 3 described above. At this time, a judgment formula was created using the product of BRCA2's SYBR-P and CDH13's SYBR-P as the test variable, and this was designated as index C (machine learning model C). Index C = 0.7716 + (-2.0705 × SYBR-P of BRCA2) + (-0.7953 × SYBR-P of CDH13) + (3.9495 × SYBR-P of BRCA2 × SYBR-P of CDH13) ROC analysis was performed using index C, and the cutoff value (0.471169) was determined from Fisher's exact test. The performance of index C was as follows: AUC = 0.852, sensitivity = 95.0%, specificity = 65.0%, positive predictive value (PPV) = 73.1%, and negative predictive value (NPV) = 92.9%.

[0077] To evaluate the reproducibility of detecting the low-risk group using Indicator C, we validated Indicator C using a cohort consisting of test data (10 cases from the low-risk group and 10 cases from the high-risk group). The diagnostic performance was AUC=0.820, sensitivity=90.0%, specificity=70.0%, PPV=75.0%, and NPV=87.5%. Therefore, we confirmed that Indicator C can reproducibly detect the low-risk group. Furthermore, Indicator C also showed high diagnostic performance in detecting malignant cases in the test data.

[0078] [Indicator D] Aside from changing the number of training data points to 40 and the number of test data points to 20, the same procedure as for metric C was used to create the judgment formula, which was designated as metric D (machine learning model D).

[0079] Index D = 0.7763 + (-2.1824 × SYBR-P of BRCA2) + (-0.7678 × SYBR-P of CDH13) + (3.7122 × SYBR-P of BRCA2 × SYBR-P of CDH13) ROC analysis was performed using index D, and the cutoff value (0.497028) was determined from Fisher's exact test. The performance of index D was as follows: AUC = 0.836, sensitivity = 93.3%, specificity = 70.0%, positive predictive value (PPV) = 75.7%, and negative predictive value (NPV) = 91.3%.

[0080] 〔result〕 [Comparison of Indicator B and Cytological Examination] Samples from 17 patients diagnosed with SCC and 6 patients diagnosed with CIS were evaluated using cytology and Index B, methods commonly used for diagnosing oral cancer. Mucosal cells used for cytology were collected by directly brushing the lesions in the oral cavity after collecting mouthwash. The cells were floated in liquid (liquid cytology), stained with Papanicolaou, and evaluated using the Bethesda System.

[0081] Cytological examination revealed that 11 out of 17 SCC patients and 1 out of 6 CIS patients were classified as having low susceptibility, for a total of 12 cases (52.2%). On the other hand, using index B, 15 out of 17 SCC patients (88.2%) and 5 out of 6 CIS patients (83.3%) were classified as having low susceptibility, for a total of 20 cases (87.0%). Therefore, it was demonstrated that the diagnostic method according to one embodiment of the present invention can determine the low susceptibility group with higher accuracy than cytological examination, regardless of whether the patient has SCC or CIS.

[0082] [Comparison before and after surgery] Of the 30 patients classified as having a low risk of developing cancer, 14 SCC patients four weeks post-surgery were assessed using index B with oral rinse solution before and after surgery. Eleven of the 14 patients (78.6%) were classified as having a low risk of developing cancer before surgery and a high risk of developing cancer after surgery, suggesting that oral cancer cells were removed by surgery. For the remaining three patients, the risk of cancer cells being present in areas other than the surgical site in the oral cavity was suggested.

[0083] [Application to OPMD] Using index D, the likelihood of developing oral cancer was assessed for seven subjects based on their oral rinse solution. The results are shown in Tables 1-3. In Tables 1-3, "control" refers to the positive control. To standardize for differences caused by subtle variations in the amount of PCR reagent prepared, control measurements were performed each time the test was conducted.

[0084] [Table 1]

[0085] [Table 2]

[0086] [Table 3]

[0087] As shown in Tables 1-3, only Subject 1 had a value for Indicator D above the cutoff value, and was therefore determined to have a low risk of developing oral cancer.

[0088] Based on the above, the likelihood of developing oral cancer could be determined with greater accuracy than cytology for all indicators A to D. Furthermore, because real-time PCR is employed, the results can be obtained quickly with simple procedures. [Industrial applicability]

[0089] This invention can be used as a method for determining the likelihood of developing oral cancer. [Explanation of symbols]

[0090] 1 Control Unit 2 Storage section 10 Information Processing Devices 11 Input section 12 Acquisition Department 13 Judgment section 14 Output Control Unit 15 Output section 21 Machine Learning Models 22 Target SYBR ratio 23 Judgment results

Claims

1. A method for determining the likelihood of developing oral cancer, which is performed by an information processing device, Extracted from the target oral rinse solution, Gene A: A gene that is methylated infrequently, regardless of whether or not one has oral cancer. Gene B: A gene in which the frequency of methylation is more than twice as high in patients with oral cancer compared to patients without oral cancer. Gene C: A gene that is frequently methylated, regardless of whether or not the person has oral cancer. The acquisition step involves obtaining the first SYBR ratio, which is the ratio of the first SYBR value to the second SYBR value, and the second SYBR ratio, which is the ratio of the first SYBR value to the third SYBR value, from the SYBR values ​​of genes A to C, namely the first SYBR value, the second SYBR value, and the third SYBR value. The method includes: (1) an explanatory variable, including the first SYBR ratio and the second SYBR ratio, calculated using the respective SYBR values ​​of genes A to C extracted from each mouthwash of a first specimen suffering from oral cancer, to which an objective variable indicating susceptibility to oral cancer is associated; and (2) an explanatory variable, including the first SYBR ratio and the second SYBR ratio, calculated using the respective SYBR values ​​of genes A to C extracted from each mouthwash of a second specimen not suffering from oral cancer, to which an objective variable indicating susceptibility to oral cancer is associated; and a determination step of determining the level of susceptibility to oral cancer of the subject based on the first SYBR ratio and the second SYBR ratio of the subject, using a machine learning model generated by machine learning with training data, wherein the explanatory variable, including the first SYBR ratio and the second SYBR ratio, calculated using the respective SYBR values ​​of genes A to C extracted from each mouthwash of a second specimen not suffering from oral cancer, to which an objective variable indicating susceptibility to oral cancer is associated; A method for determining the likelihood of developing oral cancer.

2. The method according to claim 1, wherein the first SYBR ratio and the second SYBR ratio are, respectively, first SYBR-P and second SYBR-P.

3. In the acquisition step described above, the product of the target first SYBR-P and the target second SYBR-P is further acquired. In the determination step, the explanatory variables of the machine learning model further include the product of the first SYBR-P and the second SYBR-P for each of the first and second samples, and the product of the first SYBR-P and the second SYBR-P of the target is further used to determine the degree of the target's susceptibility to oral cancer. The method according to claim 2.

4. The method according to claim 1, wherein the subject and the specimen are of the same biological species.

5. The method according to claim 1, wherein gene A is the MLH1 gene, gene B is the BRCA2 gene, and gene C is the CDH13 gene.

6. For each of the first specimen, which was affected by oral cancer, and the second specimen, which was not affected by oral cancer, the following was extracted from the mouthwash: Gene A: A gene that is methylated infrequently, regardless of whether or not one has oral cancer. Gene B: A gene in which the frequency of methylation is more than twice as high in patients with oral cancer compared to patients without oral cancer. Gene C: A gene that is frequently methylated, regardless of whether or not the person has oral cancer. The learning data acquisition step involves obtaining the first SYBR ratio, which is the ratio of the first SYBR value to the second SYBR value, and the second SYBR ratio, which is the ratio of the first SYBR value to the third SYBR value, from the SYBR values ​​of genes A to C, namely the first SYBR value, the second SYBR value, and the third SYBR value. The method includes a generation step of generating a machine learning model by performing machine learning using training data in which: (1) explanatory variables, including the first SYBR ratio and the second SYBR ratio, calculated using the respective SYBR values ​​of genes A to C extracted from each mouthwash of a first specimen affected by oral cancer, are associated with an objective variable indicating a low susceptibility to oral cancer; and (2) explanatory variables, including the first SYBR ratio and the second SYBR ratio, calculated using the respective SYBR values ​​of genes A to C extracted from each mouthwash of a second specimen not affected by oral cancer, are associated with an objective variable indicating a high susceptibility to oral cancer. How to create a machine learning model.

7. A method for determining the likelihood of developing oral cancer, which is performed by an information processing device, Extracted from the target oral rinse solution Gene A: A gene that is methylated infrequently, regardless of whether or not one has oral cancer. Gene B: A gene in which the frequency of methylation is more than twice as high in patients with oral cancer compared to patients without oral cancer. Gene C: A gene that is frequently methylated, regardless of whether or not the person has oral cancer. Regarding this, the acquisition unit acquires the first SYBR ratio, which is the ratio of the first SYBR value to the second SYBR value, and the second SYBR ratio, which is the ratio of the first SYBR value to the third SYBR value, from the SYBR values ​​of genes A to C, which are the first SYBR value, second SYBR value, and third SYBR value, respectively. The system includes: (1) an explanatory variable, including the first SYBR ratio and the second SYBR ratio, calculated using the respective SYBR values ​​of genes A to C extracted from the mouthwash of each first sample affected by oral cancer, to which an objective variable indicating a low susceptibility to oral cancer is associated; and (2) an explanatory variable, including the first SYBR ratio and the second SYBR ratio, calculated using the respective SYBR values ​​of genes A to C extracted from the mouthwash of each second sample not affected by oral cancer, to which an objective variable indicating a high susceptibility to oral cancer is associated; and a determination unit that determines the level of susceptibility to oral cancer of the subject based on the first SYBR ratio and the second SYBR ratio of the subject, using a machine learning model generated by machine learning with training data. An information processing device for determining the likelihood of developing oral cancer.

8. A control program for causing a computer to function as an information processing device according to claim 7, wherein the computer functions as the acquisition unit and the determination unit.

9. A computer-readable recording medium that stores the control program described in claim 8.

Citation Information

Patent Citations

  • Detection method of oral cavity precancerous lesion

    JP2017046628A