Tumor specific gene mutation and methylation combined bioinformatics analysis method and system

Through the combination of tumor-specific gene mutation and methylation combined with bioinformatics analysis methods, combined with mutation site data and methylation data, the limitations of detecting mutations or methylation status in the prior art are solved, achieving higher detection accuracy and accuracy in cancer type judgment.

CN120048346AActive Publication Date: 2025-05-27ZHONGKE JINCHEN BIOTECHNOLOGY (HEFEI) CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510113667.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-27
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

In the prior art, gene detection methods and analytical methods mostly focus on detecting mutations alone or detecting methylation status alone, which has limitations, especially in the early stage of cancer detection, with low accuracy.

Method used

The bioinformatics analysis method of tumor-specific gene mutation and methylation is used to improve the accuracy of the detection results through the joint analysis of mutation site data and methylation data, and multiple detections and multiple models are used to jointly judge.

Benefits of technology

By combining mutation site data and methylation data, the accuracy of prediction of sample type can be effectively improved, accurately judge whether the sample is a cancer type, and output the corresponding cancer type.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048346A_ABST
    Figure CN120048346A_ABST
Patent Text Reader

Abstract

The invention provides a tumor specific gene mutation and methylation combined bioinformatics analysis method and a tumor specific gene mutation and methylation combined bioinformatics analysis system, which are applied to the technical field of bioinformatics analysis and comprise the following steps: acquiring data of each mutation site and methylation data of a sample; judging for the first time according to the overall omega value of the sample; if the overall omega value is greater than 1, the cancer patient type is determined; if not, the Zscore value is judged for the second time; if the judgment condition I is met, the cancer patient type is judged; if not, the Zscore value is judged for the third time; if the judgment condition II is met, judging the type as a healthy type; if not, performing fourth judgment according to the random forest auxiliary model; and if the random forest auxiliary model judges that the probability of the sample belonging to the healthy sample is smaller than a threshold value, performing fifth judgment according to the neural network model. Through conjoint analysis of mutation site data and methylation data, multiple times of detection are carried out in sequence, and the accuracy of a detection result is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of bioinformatics analysis, and particularly relates to a method and system for combined bioinformatics analysis of tumor-specific gene mutations and methylation. Background Art

[0002] Cancer is a major public health problem, and early detection of cancer is crucial for improving survival rates and reducing the disease burden. Liquid biopsy is a non-invasive diagnostic method that can detect circulating tumor cells or cancer-derived nucleic acids in blood or urine. Tissue biopsy is the gold standard for cancer diagnosis, but it is invasive, expensive, and can only provide information about the sampled site. In contrast, liquid biopsy is a less invasive method that analyzes blood or other body fluids to look for signs of cancer cells or genetic material. The potential of cell-free DNA (cfDNA) liquid biopsy in early detection of cancer signals and prediction of disease progression has been widely studied. Cancer genome analysis using next-generation sequencing (NGS) technology involves sequencing and analyzing the DNA of cancer cells to identify gene mutations in areas known as hotspots (such as KRAS and PIK3CA), which are associated with the occurrence and development of cancer. There are several different methods for cancer genome analysis using NGS, such as whole-genome sequencing and genomic analysis.

[0003] Using a targeted sequencing panel for hotspot mutation analysis mainly detects specific mutations common in cancer, but this method may have limitations in early cancer detection. To further overcome the limitations of traditional liquid biopsy detection based on mutations and fragments, researchers have been exploring the use of cancer-specific methylation signals. This method involves analyzing the patterns of DNA methylation changes specific to cancer cells and using these characteristics to detect the presence of cancer-specific methylation signals in cfDNA.

[0004] In addition, machine learning is a common research hotspot in the fields of artificial intelligence and pattern recognition, and its theories and methods have been widely applied to solve complex problems in engineering applications and scientific fields. With the introduction and application of machine learning in the field of computational biology, a diagnostic model is constructed based on the characteristics of cfDNA methylation levels to detect and locate potential tumors, overcoming the problem of insufficient resolution in conventional plasma diagnosis.

[0005] Using machine learning-related methods for modeling analysis is a highly feasible approach for early cancer diagnosis based on cfDNA. In June 2024, a joint research team from Imperial College London and the University of Cambridge trained an artificial intelligence model - EMethylNET, which trained and evaluated four different model types: logistic regression, support vector machine (SVM), extreme gradient boosting (XGBoost), and deep neural network (DNN). For the first three models, binary classification and multi-classification models were created respectively. By analyzing DNA methylation patterns, the model was able to identify 13 different types of cancers (including breast cancer, liver cancer, lung cancer, and prostate cancer, etc.) from non-cancerous tissues. In 2022, a team from the Second General Hospital of Guangdong Province and Tsinghua University established a semi-reference-free deconvolution (SRFD) algorithm based on machine learning methods, which can use cfDNA methylation profiles to extract tumor information and locate the origin tissue of the tumor. The research team finally established a Bayesian diagnostic model (SRFD-Bayes), and the average localization accuracy of this model for normal controls and all early tumors was 76.9%.

[0006] In summary, the gene detection methods and analysis methods in the prior art mostly focus on detecting mutations alone or detecting methylation status alone. For the method of detecting mutations alone, its liquid biopsy detection based on mutations and fragments has certain limitations and is not accurate enough in early cancer detection. For the method of detecting methylation status alone, the methylation patterns may vary among different cancer subtypes and individuals, and the methylation level gap between some cancer types and healthy people is not large, which makes it difficult to accurately distinguish cancer and healthy samples. In addition, methylation changes may be affected by factors other than cancer, leading to potential false positives. Summary of the Invention

[0007] In view of the above problems in the prior art, the purpose of the present invention is to provide a combined bioinformatics analysis method for tumor-specific gene mutations and methylation, which jointly analyzes mutation site data and methylation data and conducts multiple detections in sequence to ensure the accuracy of the detection results.

[0008] A combined bioinformatics analysis method for tumor-specific gene mutations and methylation includes the following steps:

[0009] Obtain the mutation site data and methylation data of the sample to be tested;

[0010] Based on the reference threshold of the mutation site detection model, calculate the Ω value of each mutation site in the sample to be tested and make a first judgment according to the overall Ω value;

[0011] If the overall Ω value is greater than 1, it is determined that the type of the sample to be tested is the cancerous type; if the overall Ω value is not greater than 1, a second judgment is made on the Zscore value of the methylation data based on Judgment Condition 1.

[0012] If it meets Judgment Condition 1, it is determined that the sample to be tested is of the cancerous type; if it does not meet, a third judgment is made on the Zscore value of the methylation data based on Judgment Condition 2.

[0013] If it meets Judgment Condition 2, it is determined that the sample to be tested is of the healthy type; if it does not meet, a fourth judgment is made according to the random forest auxiliary model constructed based on the methylation data.

[0014] If the probability that the random forest auxiliary model constructed based on the methylation data determines that the sample to be tested belongs to a healthy sample is less than the threshold, a fifth judgment is made according to the neural network model constructed based on the methylation data to determine whether the sample to be tested is of the cancerous type.

[0015] After determining that the type of the sample to be tested is the cancerous type, a judgment on the cancer type is made according to the random forest model for cancer type prediction to output the probability values that the sample to be tested belongs to different cancer types.

[0016] Preferably, a corresponding mutation site detection model is constructed by setting a training set and integrating the mutation frequencies of the data corresponding to the mutation sites in the training set. The training set includes cancerous samples and healthy samples. The mutation frequency of each mutation site in the mutation site detection model has a corresponding reference threshold, and the range of the reference threshold is from 0 to 0.02.

[0017] Preferably, the determination process of the reference threshold for the mutation frequency of each mutation site is as follows:

[0018] Each mutation site of each sample in the training set is processed in a loop, and the training set is divided into corresponding cancerous sample mutation frequency data sets and healthy sample mutation frequency data sets according to the mutation sites.

[0019] A reference mutation frequency is set and the proportion of samples with a mutation frequency greater than the reference mutation frequency in the healthy sample mutation frequency data set corresponding to each mutation site is calculated. Whether each mutation site is credible is judged according to the proportion value of the samples.

[0020] On the basis that the mutation site is credible, it is judged whether there is a difference in this mutation site between the cancerous sample mutation frequency data set and the healthy sample mutation frequency data set through a T-test.

[0021] In the case of no difference: The reference threshold is determined according to the maximum value of the mutation frequency in the healthy sample mutation frequency data set and the 0.99 quantile.

[0022] In the case of differences or inability to determine differences: If the maximum value of the mutation frequencies in the healthy sample mutation frequency dataset is not greater than the baseline mutation frequency, the reference threshold for this mutation site is set to the maximum value of the mutation frequencies of healthy samples in the training set; if the maximum value of the mutation frequencies in the healthy sample mutation frequency dataset is greater than the baseline mutation frequency, the reference threshold for this mutation site is set to the baseline mutation frequency.

[0023] Preferably, in the case where there are no differences in the corresponding mutation sites between the cancer sample mutation frequency dataset and the healthy sample mutation frequency dataset, the process for determining the reference threshold for this mutation site is as follows:

[0024] If the maximum value of the mutation frequencies in the healthy sample mutation frequency dataset for this mutation site is greater than or equal to 0.40, the reference threshold for this mutation site is set to 0.01;

[0025] If the maximum value of the mutation frequencies in the healthy sample mutation frequency dataset for this mutation site is less than 0.01 and 1.30 times the 0.99 - quantile is also less than 0.01, the reference threshold for this mutation site is set to 0.01;

[0026] If the maximum value of the mutation frequencies in the healthy sample mutation frequency dataset for this mutation site is greater than 0.01 and less than 0.40, and at the same time 1.30 times the 0.99 - quantile is also greater than 0.01, then, if the maximum value is greater than 1.30 times the 0.99 - quantile, the reference threshold for this mutation site is set to the maximum value of the mutation frequencies of healthy samples in the training set; if the maximum value is less than or equal to 1.30 times the 0.99 - quantile, the reference threshold for this mutation site is set to 1.30 times the 0.99 - quantile of the mutation frequencies of healthy samples in the training set;

[0027] If the maximum value of the mutation frequencies in the healthy sample mutation frequency dataset for this mutation site is greater than 0.01 and less than 0.40, and at the same time 1.30 times the 0.99 - quantile is less than or equal to 0.01, the reference threshold for this mutation site is set to the maximum value of the mutation frequencies of healthy samples in the training set;

[0028] If the maximum value of the mutation frequencies in the healthy sample mutation frequency dataset for this mutation site is less than or equal to 0.01 and 1.30 times the 0.99 - quantile is greater than 0.01, the reference threshold for this mutation site is set to 1.30 times the 0.99 - quantile of the mutation frequencies of healthy samples in the training set.

[0029] Preferably, the Ω value represents the degree of proximity between the mutation frequency of a mutation site in a test sample and that of a cancer sample. The overall Ω value of the test sample is the sum of the Ω values of each mutation site in the test sample. The process for calculating the Ω value of each mutation site in the test sample is as follows:

[0030] If the mutation frequency of the mutation site in the sample to be tested is greater than or equal to 0.40, then let Ω = 1;

[0031] If the mutation frequency of the mutation site in the sample to be tested is less than or equal to the reference threshold, if the reference threshold is less than 0.01, then let Ω = 99.999 * AF; if the reference threshold is greater than or equal to 0.01, then let Ω = 50 * AF;

[0032] If the mutation frequency of the mutation site in the sample to be tested is greater than the reference threshold, if the reference threshold is less than 0.01 and the mutation frequency of the sample to be tested is greater than 0.005, then let Ω = 99.999 * AF + 1; if the reference threshold is less than 0.01 and the mutation frequency of the sample to be tested is less than or equal to 0.005, then let Ω = 99.999 * AF; if the reference threshold is greater than or equal to 0.01 and the mutation frequency of the sample to be tested is greater than 0.005, then let Ω = 50 * AF + 1; if the reference threshold is greater than or equal to 0.01 and the mutation frequency of the sample to be tested is less than or equal to 0.005, then let Ω = 50 * AF;

[0033] Wherein, the AF value is used to represent the frequency of each allele at a certain mutation site; the AF value is a numerical value between 0 and 1, where 0 means the frequency of this allele among all alleles is 0, and 1 means the frequency of this allele among all alleles is 1.

[0034] Preferably, the first judgment condition is whether the proportion of Zscore values greater than 3 in the methylation data is greater than 7%, or whether the maximum value of the Zscore value is greater than 10, or whether the average value of the Zscore value is greater than 0.4;

[0035] The second judgment condition is whether the proportion of Zscore values greater than 1 in the methylation data is less than 3%, or whether the maximum value of the Zscore value is less than 1.80, or whether the average value of the Zscore value is less than -1.05.

[0036] Preferably, in the random forest auxiliary model constructed based on methylation data, if the probability that the sample to be tested belongs to a healthy sample is not less than 65%, then it is determined that the type of the sample to be tested is a healthy type; if the probability that the sample to be tested belongs to a healthy sample is less than 65%, then the fifth judgment is made according to the neural network model constructed based on methylation data.

[0037] Preferably, in the neural network model constructed based on methylation data, if the probability value that the sample to be tested belongs to a healthy sample is greater than 50%, then it is determined that the type of the sample to be tested is a healthy type, if the probability value that the sample to be tested belongs to a healthy sample is not greater than 50%, then it is determined that the type of the sample to be tested is a cancerous type.

[0038] Preferably, the process of calculating the Zscore value of the methylation data of the sample to be tested is as follows:

[0039] Set a training set, which includes cancer samples and healthy samples. The cancer samples include samples of different types of cancer patients.

[0040] Calculate the average value μ and standard deviation σ of each methylation site of the healthy samples in the training set.

[0041] Calculate the Zscore value of the sample to be tested compared to the training set. The calculation formula is: where X represents the methylation level value of the corresponding methylation site of the sample to be tested.

[0042] The second object of the present invention is to propose a tumor-specific gene mutation and methylation combined bioinformatics analysis system, including a data processing module, a mutation site detection module, a methylation detection module, and a cancer type detection module. The methylation detection module includes a Zscore value detection model, a random forest auxiliary model, and a neural network model. The data processing module is used to perform upstream analysis on the sample to be tested to obtain the data of each mutation site and the methylation data of the sample. The mutation site detection module is used to calculate the Ω value of each mutation site in the sample to be tested based on the reference threshold of the mutation site detection model and output whether the sample to be tested is a cancer type according to the overall Ω value. The methylation detection module is used to sequentially judge the methylation data of the sample to be tested through the Zscore value detection model, the random forest auxiliary model, and the neural network model, and output whether the sample to be tested is a cancer type. The cancer type detection module is used to judge which cancer type the cancerous sample to be tested belongs to based on the methylation data when the mutation site detection module or the methylation detection module outputs that the sample to be tested is a cancer type.

[0043] The beneficial effects of the present invention are as follows: For the tumor-specific gene mutation and methylation combined bioinformatics analysis method and system, first, the cancer probability is detected by the mutation site detection module based on the mutation site data. After detecting that the sample is not a cancer type, the methylation data is continuously detected by multiple models constructed based on machine learning. By combining the mutation site data and the methylation data and using the method of joint judgment by multiple models, the accuracy of sample type prediction can be effectively improved. And after judging that the sample type is a cancer type, the cancer type can be further judged by the analysis model, and the corresponding cancer type can be output.

[0044] When performing upstream analysis on the sample, the mutation site data and methylation status can be analyzed simultaneously. Compared with the method of performing mutation detection or methylation detection alone, the processes of sequence filtering and aligning to the reference genome in the analysis process of this analysis method are exactly the same, which can save analysis time and computing resources. Description of the Drawings

[0045] The accompanying drawings are used to provide a further understanding of the present invention and form a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the accompanying drawings:

[0046] Figure 1 is the system flowchart of the present invention;

[0047] Figure 2 is the flowchart of the upstream analysis of the mutation site data of the present invention;

[0048] Figure 3 is the flowchart of the upstream analysis of the methylation data of the present invention;

[0049] Figure 4 is a schematic diagram of the correct prediction ratio of the top three types with the highest prediction probabilities in the male sample cancer type prediction model constructed based on the present invention;

[0050] Figure 5 is a schematic diagram of the correct prediction ratio of the top three types with the highest prediction probabilities in the female sample cancer type prediction model constructed based on the present invention. Detailed implementation manners

[0051] Example 1

[0052] As Figure 1 shown, a combined bioinformatics analysis method for tumor-specific gene mutations and methylation specifically includes the following steps:

[0053] First step, perform upstream analysis on the sample to be tested to obtain the mutation site data and methylation data of each sample.

[0054] Among them, the mutation site data includes the mutation frequency, sequencing depth, and annotation results of the corresponding mutation sites. The methylation data includes the methylation level, and the calculation formula is:

[0055]

[0056] In addition, the results of the upstream analysis of the sample to be tested also include the QC data of the mutation sites and the QC data of methylation. The QC data of the mutation sites includes the alignment rate, sequencing depth, coverage, and sequencing uniformity. The QC data of methylation includes the digestion efficiency and sequencing depth. The digestion efficiency is used to evaluate the cutting effect of the restriction endonuclease sensitive to methylation. A poor cutting effect will lead to inaccurate calculation of the methylation level and affect the reliability of subsequent analysis. Generally, it should be greater than 98%, and it is calculated based on the sequencing depth of the exogenous DNA (Lambda DNA) added in the experiment. The calculation formula is:

[0057]

[0058] Specifically, as Figure 2 shown, the process of obtaining mutant site data is as follows:

[0059] (1) Use the software trim_galore to filter adapters and low-quality sequences in the raw data.

[0060] (2) Use the software BWA to align the filtered sequences to the human reference genome (GRCh37 / hg19), and output the result as a BAM file.

[0061] (3) Use MarkDuplicates in the GATK tool to remove PCR duplicates in the sequences.

[0062] (4) Use the software Vardict to detect mutant sites in the sequences. A mutant site refers to a change in a single base nucleotide in the DNA or RNA sequence, including transitions, transversions, insertions, or deletions.

[0063] (5) Use the software ANNOVAR and the annotation information database to annotate the mutant sites.

[0064] (6) Obtain mutant site data.

[0065] As Figure 3 shown, the process of obtaining methylation data is as follows:

[0066] (1) Use the software trim_galore to filter low-quality sequences and remove adapters in the raw data.

[0067] (2) Use the software BWA to align the filtered sequences to the human reference genome (GRCh37 / hg19), and output the result as a BAM file.

[0068] (3) Use pysam to process the BAM file, filter sequences that are not aligned to the reference genome, secondary alignment or multiple alignment sequences, and sequences with an R1 and R2 (read1 and read2 sequences in paired-end sequencing, i.e., two sequences with opposite directions in two complementary strands) interval > 2000bp, remove PCR duplicates, and then filter sequences in the adapter region.

[0069] (4) Calculate the sequencing depth of the target region and the internal reference region and calculate the methylation level.

[0070] (5) Calculate the digestion efficiency according to the sequencing depth of the added exogenous DNA (Lambda DNA).

[0071] Step 2: Based on the reference threshold of the mutation site detection model, calculate the Ω value of each mutation site in the test sample, and make a first judgment based on the overall Ω value of the test sample to determine whether the test sample belongs to the cancer type.

[0072] The process of the first judgment: If the overall Ω value of the test sample is greater than 1, it is determined that the type of the test sample is the cancer type; if the overall Ω value of the test sample is not greater than 1, a second judgment is made on the Zscore value of the methylation data based on judgment condition one.

[0073] Specifically, a corresponding mutation site detection model is constructed by setting a training set and integrating the mutation frequencies of the corresponding mutation site data in the training set. The training set includes cancer samples and healthy samples. It should be noted that male samples and female samples are detected separately.

[0074] Among them, each mutation site in the mutation site detection model has a corresponding reference threshold. The determination process of the reference thresholds of the mutation frequencies of each mutation site is as follows:

[0075] Loop through each mutation site of each sample in the training set, and divide the training set into corresponding cancer sample mutation data sets and healthy sample mutation data sets according to the mutation sites.

[0076] Set the baseline mutation frequency as A and calculate the proportion of samples in the healthy sample mutation frequency data set corresponding to each mutation site whose mutation frequency is greater than A. Judge whether each mutation site is credible according to the sample proportion value. Specifically, set the baseline mutation frequency as 0.01. For each mutation site, calculate the proportion of samples in the healthy sample mutation frequency data set whose mutation frequency is greater than 0.01. If the proportion of samples in the healthy sample mutation frequency data set corresponding to this mutation site whose mutation frequency is greater than 0.01 is greater than 1%, it is considered that this mutation site is not credible, and all data of this mutation site will be deleted in subsequent analysis.

[0077] If the number of cancer samples and healthy samples detected for this mutation site in the training set is greater than 1, perform a T - test on the cancer sample mutation frequency data set and the healthy sample mutation frequency data set. The T - test is a statistical method mainly used to evaluate whether there is a significant difference in the means of two samples. If the P - value of the T - test is greater than 0.05, it is considered that there is no difference in this mutation site between cancer samples and healthy samples; if the P - value of the T - test is less than 0.05, it is considered that there is a difference in this mutation site between cancer samples and healthy samples.

[0078] In the case where there is no difference in the corresponding mutation sites between the cancer sample mutation data set and the healthy sample mutation data set:

[0079] If the maximum mutation frequency in the mutation frequency dataset of healthy samples at this mutation site is greater than or equal to 0.40, the reference threshold for this mutation site is set to 0.01.

[0080] If the maximum mutation frequency in the mutation frequency dataset of healthy samples at this mutation site is less than 0.01 and 1.30 times the 0.99th percentile (in a set of data, 99% of the data is less than or equal to this value) is also less than 0.01, the reference threshold for this mutation site is set to 0.01.

[0081] If the maximum mutation frequency in the mutation frequency dataset of healthy samples at this mutation site is greater than 0.01 and less than 0.40, and at the same time 1.30 times the 0.99th percentile is also greater than 0.01. At this time, if the maximum value is greater than 1.30 times the 0.99th percentile, the reference threshold for this mutation site is set to the maximum mutation frequency of healthy samples in the training set; if the maximum value is less than or equal to 1.30 times the 0.99th percentile, the reference threshold for this mutation site is set to 1.30 times the 0.99th percentile of the mutation frequency of healthy samples in the training set.

[0082] If the maximum mutation frequency in the mutation frequency dataset of healthy samples at this mutation site is greater than 0.01 and less than 0.40, and 1.30 times the 0.99th percentile is less than or equal to 0.01, the reference threshold for this mutation site is set to the maximum mutation frequency of healthy samples in the training set.

[0083] If the maximum mutation frequency in the mutation frequency dataset of healthy samples at this mutation site is less than or equal to 0.01 and 1.30 times the 0.99th percentile is greater than 0.01, the reference threshold for this mutation site is set to 1.30 times the 0.99th percentile of the mutation frequency of healthy samples in the training set.

[0084] In the case where there are differences or the difference cannot be determined at the corresponding mutation sites in the cancer sample mutation dataset and the healthy sample mutation dataset:

[0085] If the maximum mutation frequency in the mutation frequency dataset of healthy samples at this mutation site is less than or equal to 0.01, the reference threshold for this mutation site is set to the maximum mutation frequency of healthy samples in the training set; if the maximum mutation frequency in the mutation frequency dataset of healthy samples at this mutation site is greater than 0.01, the reference threshold for this mutation site is set to 0.01.

[0086] In the process of determining the reference threshold for the mutation frequency of each mutation site in the mutation site detection model, the values used for judgment are determined after multiple adjustments and tests in previous studies.

[0087] The Ω value represents the proximity of the mutation frequency at the mutation site in the sample to be tested to that in the cancerous sample. The overall Ω value of the sample to be tested is the sum of the Ω values of each mutation site in the sample to be tested. The process of calculating the Ω value of each mutation site in the sample to be tested is as follows:

[0088] If the mutation frequency at the mutation site in the sample to be tested is greater than or equal to 0.40, then let Ω = 1.

[0089] If the mutation frequency at the mutation site in the sample to be tested is less than or equal to the reference threshold, and if the reference threshold is less than 0.01, then let Ω = 99.999 * AF; if the reference threshold is greater than or equal to 0.01, then let Ω = 50 * AF.

[0090] If the mutation frequency at the mutation site in the sample to be tested is greater than the reference threshold, and if the reference threshold is less than 0.01 and the mutation frequency in the sample to be tested is greater than 0.005, then let Ω = 99.999 * AF + 1; if the reference threshold is less than 0.01 and the mutation frequency in the sample to be tested is less than or equal to 0.005, then let Ω = 99.999 * AF; if the reference threshold is greater than or equal to 0.01 and the mutation frequency in the sample to be tested is greater than 0.005, then let Ω = 50 * AF + 1; if the reference threshold is greater than or equal to 0.01 and the mutation frequency in the sample to be tested is less than or equal to 0.005, then let Ω = 50 * AF.

[0091] Among them, the AF value is used to represent the frequency of each allele at a certain mutation site. The AF value is a numerical value between 0 and 1, where 0 means the frequency of this allele among all alleles is 0, and 1 means the frequency of this allele among all alleles is 1.

[0092] Step 3: Based on Judgment Condition 1, make a second judgment on the Zscore value of the methylation data.

[0093] If the Zscore value of the methylation data meets Judgment Condition 1, then determine that the type of the sample to be tested is the cancerous type; if the Zscore value of the methylation data does not meet Judgment Condition 1, then based on Judgment Condition 2, make a third judgment.

[0094] Judgment Condition 1: Whether the proportion of the Zscore value of the methylation data greater than 3 is greater than 7%, or whether the maximum value of the Zscore value is greater than 10, or whether the average value of the Zscore value is greater than 0.4.

[0095] Step 4: Based on Judgment Condition 2, make a third judgment on the Zscore value of the methylation data.

[0096] If the Zscore value of the methylation data meets the second judgment condition, the type of the sample to be tested is determined as the healthy type; if the Zscore value of the methylation data does not meet the second judgment condition, the fourth judgment is made according to the random forest auxiliary model constructed based on the methylation data.

[0097] The second judgment condition: whether the proportion of Zscore values greater than 1 in the methylation data is less than 3%, or whether the maximum value of the Zscore value is less than 1.80, or whether the average value of the Zscore value is less than -1.05.

[0098] Among them, the process of calculating the Zscore value of the methylation data of the sample to be tested is as follows:

[0099] Set the training set. The training set includes cancer samples and healthy samples, and the cancer samples include samples of different types of cancer patients. It should be noted that male samples and female samples are detected separately.

[0100] Calculate the average value μ and standard deviation σ of each methylation site of the healthy samples in the training set.

[0101] Calculate the Zscore value of the sample to be tested compared with the training set, and the calculation formula is: Among them, X represents the methylation level value of the corresponding methylation site of the sample to be tested.

[0102] In the process of judging the Zscore value of the methylation data:

[0103] When the proportion of Zscore values greater than 3 in the methylation data is greater than 7%, or the maximum value of the Zscore value is greater than 10, or the average value of the Zscore value is greater than 0.4, the type of the sample to be tested is the cancer type.

[0104] When the proportion of Zscore values greater than 1 in the methylation data is less than 3%, or the maximum value of the Zscore value is less than 1.80, or the average value of the Zscore value is less than -1.05, the type of the sample to be tested is the healthy type.

[0105] When the Zscore value of the methylation data does not belong to the above two situations, the fourth judgment is made according to the random forest auxiliary model constructed based on the methylation data.

[0106] Step 5. In the random forest auxiliary model constructed based on the methylation data, if the probability that the sample to be tested belongs to the healthy sample is not less than 65%, the type of the sample to be tested is determined as the healthy type; if the probability that the sample to be tested belongs to the healthy sample is less than 65%, the fifth judgment is made according to the neural network model constructed based on the methylation data.

[0107] Specifically, the process of constructing a random forest auxiliary model based on methylation data is as follows:

[0108] Set up the training set. The training set includes cancer samples and healthy samples, and the cancer samples include samples from different types of cancer patients. It should be noted that male samples and female samples are detected separately.

[0109] Construct a random forest auxiliary model and train the random forest auxiliary model according to the training set. The random forest auxiliary model includes multiple decision trees, and each decision tree extracts corresponding samples from the training set by the method of replacement extraction.

[0110] The process of constructing a random forest auxiliary model based on methylation data is the same as that of constructing a random forest model in the prior art, that is, first, by sampling with replacement from the training set for several times to form a new sub-training set D, then randomly selecting m features, and then using the new training set D and m features to learn a complete decision tree. Repeat the above steps K times to obtain K decision trees and form a random forest.

[0111] Step 6: In the neural network model constructed based on methylation data, if the probability value that the sample to be tested belongs to a healthy sample is greater than 50%, it is determined that the type of the sample to be tested is a healthy type; if the probability value that the sample to be tested belongs to a healthy sample is not greater than 50%, it is determined that the type of the sample to be tested is a cancer type.

[0112] A neural network model is a machine learning model composed of multiple neuron layers. Each neuron layer receives the output of the previous layer as input and calculates the output through a series of non-linear transformations and weight adjustments. Then it is trained by the backpropagation algorithm, by calculating the error between the predicted output and the true output, and using the gradient descent method to update the weights and bias values in the network until the network reaches the predetermined performance level.

[0113] The neural network model in this embodiment is trained and constituted by a training set of methylation level data including cancer samples and healthy samples. Its input layer is the methylation levels corresponding to 431 methylation sites; the output layer is 0 (representing healthy type samples) and 1 (representing cancer type samples); the backpropagation algorithm is used to update the weight values; the activation function used in the hidden layer of the neural network model is the rectified linear activation function; the activation function used in the output layer is the softmax function.

[0114] Step 7: After determining that the type of the sample to be tested is a cancer type, judge the cancer type according to the random forest model for cancer type prediction to output the probability values that the sample to be tested belongs to different cancer types.

[0115] The output of the random forest model for cancer type prediction is divided into healthy types and specific cancer types, including bladder cancer, bile duct cancer, colorectal cancer, esophageal cancer, kidney cancer, liver cancer, lung cancer, pancreatic cancer, gastric cancer, and thyroid cancer.

[0116] In this embodiment, the parameters of the neural network model constructed by methylation data, the random forest auxiliary model constructed based on methylation data, and the random forest model for cancer type prediction are optimized by the ten-fold cross validation method. Specifically, the ten-fold cross validation method divides the data set into 10 parts, and 9 of them are used as training data and 1 is used as test data for testing in turn. Each test will obtain the corresponding accuracy (or error rate), and then the average accuracy (or error rate) of the 10 results is used as an estimate of the accuracy of the algorithm.

[0117] Embodiment 2

[0118] This embodiment proposes a tumor-specific gene mutation and methylation joint bioinformatics analysis system, including a data processing module, a mutation site detection module, a methylation detection module and a cancer type detection module. The methylation detection module includes a Zscore value detection model, a random forest auxiliary model, and a neural network model.

[0119] The data processing module is used to perform upstream analysis on the sample to be tested to obtain data on each mutation site of the sample and methylation data of the sample.

[0120] The mutation site detection module is used to cyclically process each mutation site in the sample to be tested, compare the mutation frequency in each mutation site data with the corresponding reference threshold, and output the result of whether the sample to be tested is a cancer type.

[0121] The methylation detection module is used to judge the methylation data of the sample to be tested through the Zscore value detection model, the random forest auxiliary model, and the neural network model in sequence, and output the result of whether the sample to be tested is a cancer type.

[0122] The cancer type detection module is used to determine which cancer type the sample to be tested belongs to based on methylation data when the mutation site detection module or the methylation detection module outputs that the sample to be tested is a cancer type.

[0123] After the data processing module obtains the data of each mutation site of the sample and the methylation data of the sample, the first cancer type judgment is performed by the mutation site detection module. After the mutation site detection module fails to detect the cancer type, the Zscore value detection model performs the second cancer type judgment according to judgment condition 1. If the cancer type is not detected, the Zscore value detection model performs the third cancer type judgment according to judgment condition 2. If the cancer type is not detected, the random forest auxiliary model performs the fourth cancer type judgment. If the cancer type is not detected, the neural network model performs the fifth cancer type judgment. Among them, after the cancer type is detected, the cancer type detection module outputs which cancer type the sample to be tested belongs to.

[0124] Example 3

[0125] In this embodiment, a male analysis model and a female analysis model of the combined bioinformatics of methylation and mutation sites related to human tumors are constructed respectively.

[0126] The training set of the male analysis model includes 1 bladder cancer (BLCA), 16 colorectal cancers (COADREAD), 36 esophageal cancers (ESCA), 19 liver cancers (LIHC), 69 lung cancers (LUNG), 2 pancreatic cancers (PAAD), 18 gastric cancers (STAD), 3 thyroid cancers (THCA), and 87 healthy samples. The test set is 12 colorectal cancers (COADREAD), 15 esophageal cancers (ESCA), 8 liver cancers (LIHC), 39 lung cancers (LUNG), 14 gastric cancers (STAD), 1 thyroid cancer (THCA), and 35 healthy samples.

[0127] Based on the analysis method described in Example 1, the result statistical table 1 for judging cancerous samples or normal samples in the above male analysis model is obtained.

[0128] Table 1 Result statistical table for judging cancerous samples or normal samples in the male analysis model

[0129]

[0130] As can be seen from the data in Table 1, in the male analysis model, the specificity of the results of the training set for judging cancerous samples or normal samples is 100.00%, and the sensitivity is 99.39%; the specificity of the results of the test set for judging cancerous samples or normal samples is 91.43%, and the sensitivity is 89.89%.

[0131] The training set of the female analysis model includes 1 bladder cancer (BLCA), 3 breast cancers (BRCA), 41 cervical cancers (CESC), 15 colorectal cancers (COADREAD), 8 esophageal cancers (ESCA), 1 renal chromophobe cell carcinoma (KICH), 8 liver cancers (LIHC), 90 lung cancers (LUNG), 4 ovarian cancers (OV), 11 gastric cancers (STAD), 4 thyroid cancers (THCA), 27 endometrial cancers (UCEC), and 77 healthy samples. The test set consists of 2 breast cancers (BRCA), 17 cervical cancers (CESC), 1 cholangiocarcinoma (CHOL), 13 colorectal cancers (COADREAD), 3 esophageal cancers (ESCA), 4 liver cancers (LIHC), 40 lung cancers (LUNG), 6 ovarian cancers (OV), 12 gastric cancers (STAD), 1 thyroid cancer (THCA), 16 endometrial cancers (UCEC), and 38 healthy samples.

[0132] Based on the analysis method described in Example 1, the result statistical table 2 for judging cancerous samples or normal samples in the above female analysis model is obtained.

[0133] Table 2 Statistical table of the results for judging cancerous samples or normal samples in the female analysis model

[0134]

[0135] From the data in Table 2, it can be seen that in the female analysis model, the specificity of the results for judging cancerous samples or normal samples in the training set is 100.00%, and the sensitivity is 94.84%; the specificity of the results for judging cancerous samples or normal samples in the test set is 92.11%, and the sensitivity is 89.57%.

[0136] Among them, specificity refers to the proportion of negative test results among the population found not to have cancer, specificity = true negative population / population without cancer. Sensitivity refers to the proportion of positive test results among cancer patients, reflecting the ability to judge positive cases, sensitivity = true positive population / all cancer patients. The ideal values of both specificity and sensitivity are 100%.

[0137] According to Table 1 and Table 2, it can be seen that after the combined analysis of methylation and mutation sites, both the specificity and sensitivity of the test set are better than the results of single analysis.

[0138] As Figure 4 shown, in the cancer type prediction analysis model for men, the likelihood of correctly predicting liver cancer (100%), lung cancer (100%), esophageal cancer (93.33%), and colorectal cancer (86.67%) among the top three cancer types with predicted probabilities in the test set is greater than 80%.

[0139] AsFigure 5 As shown, in the cancer type prediction and analysis model for females, the probabilities of correctly predicting cervical cancer (100%), cholangiocarcinoma (100%), lung cancer (100%), and endometrial cancer (87.50%) in the test set among the top three cancer types in terms of prediction probability are all greater than 80%.

[0140] Example 4

[0141] In this example, Sample 1 is from a 63-year-old male patient diagnosed with stage 1 lung cancer. Based on the methylation level, the probability of it being a cancerous sample is 98.88%; in cancer type prediction, the top three cancer types in terms of probability are lung cancer (43.20%), esophageal cancer (17.00%), and gastric cancer (15.00%). The mutation site with the highest Ω value is PTEN:NM_000314.8:exon5:c.318del:p.D107Ifs*6, with a mutation frequency of 0.0052. This mutation can produce various forms of C-terminal truncated PTEN proteins. The truncating mutation near the N-terminus results in the loss of PTEN phosphatase function and the inability to negatively regulate the activity of the PI3K / AKT pathway (PMID: 11237521). Expressing the PTEN truncating mutation in mouse embryonic fibroblasts indicates that these mutations are carcinogenic. Compared with the full-length protein, PTEN cannot bind to the centrosome of the chromosome, thus increasing genomic vulnerability (PMID: 17218262).

[0142] Table 3 Detection Results of Mutation Sites in Actual Samples

[0143]

[0144] In this example, Sample 2 is from a 50-year-old female patient diagnosed with stage 1 lung cancer. Based on the methylation level, the probability of it being a cancerous sample is 95.74%; in cancer type prediction, the top three cancer types in terms of probability are lung cancer (31.20%), colorectal cancer (24.40%), and esophageal cancer (6.60%). The mutation site with the highest Ω value is CDKN2A:NM_000077.5:c.157A>G:p.Met53Val, with a mutation frequency of 0.0092. This mutation is located in the ankyrin repeat sequence of the p16 / INK4A protein. This mutation appears as a germline variation in melanoma families (PMID: 15146471). Expressing this mutation in an RB-negative osteosarcoma cell line and a melanoma cell line shows that, compared with the wild type, this mutation may be inactivated, manifested as reduced binding to CDK4 and increased cell proliferation (PMID: 20340136, 19260062).

[0145] Table 4 Detection Results of Mutation Sites in Actual Samples

[0146]

[0147] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A combined bioinformatics analysis method for tumor-specific gene mutation and methylation, characterized in that: The steps include: Obtaining mutation site data and methylation data of the samples to be tested; Based on the reference threshold of the mutation site detection model, the Ω value of each mutation site in the sample to be tested is calculated and the first judgment is made based on the overall Ω value; If the overall Ω value is greater than 1, the sample type to be tested is determined to be a cancer type; if the overall Ω value is not greater than 1, a second judgment is made on the Zscore value of the methylation data based on condition one; If the condition 1 is met, the sample to be tested is determined to be a cancer type; if not, the methylation data Zscore value is judged for the third time based on the condition 2; If the second condition is met, the sample to be tested is judged to be healthy; if not, a fourth judgment is made based on the random forest auxiliary model built based on methylation data; If the random forest auxiliary model determines that the probability that the sample to be tested is a healthy sample is less than the threshold, the neural network model built based on the methylation data will be used to make a fifth judgment to determine whether the sample to be tested is a cancerous type; After determining that the type of the sample to be tested is a cancer type, the cancer type is determined based on the random forest model for cancer type prediction to output the probability value of the sample to be tested belonging to different cancer types.

2. The method for combined bioinformatics analysis of tumor-specific gene mutation and methylation according to claim 1, characterized in that: A corresponding mutation site detection model is constructed by setting a training set and integrating the mutation frequencies of corresponding mutation site data in the training set. The training set includes cancer samples and healthy samples. The mutation frequency of each mutation site in the mutation site detection model has a corresponding reference threshold, and the range of the reference threshold is 0 to 0.

02.

3. The method for combined bioinformatics analysis of tumor-specific gene mutation and methylation according to claim 2, characterized in that: The process of determining the reference threshold of mutation frequency at each mutation site is as follows: Each mutation site of each sample in the training set is processed cyclically, and the training set is divided into a corresponding cancer sample mutation frequency dataset and a healthy sample mutation frequency dataset according to the mutation site; Set the benchmark mutation frequency and calculate the proportion of samples with a mutation frequency greater than the benchmark mutation frequency in the healthy sample mutation frequency data set corresponding to each mutation site, and judge whether each mutation site is credible based on the sample proportion value; On the basis that the mutation site is credible, a T test is used to determine whether there is a difference in the mutation site between the mutation frequency dataset of cancer samples and the mutation frequency dataset of healthy samples; In the absence of differences: the reference threshold was determined based on the maximum value of the mutation frequency in the healthy sample mutation frequency dataset and the 0.99th fraction; In the case of differences or inability to determine differences: if the maximum value of the mutation frequency in the healthy sample mutation frequency dataset is not greater than the benchmark mutation frequency, the reference threshold of the mutation site is set to the maximum value of the healthy sample mutation frequency in the training set; if the maximum value of the mutation frequency in the healthy sample mutation frequency dataset is greater than the benchmark mutation frequency, the reference threshold of the mutation site is set to the benchmark mutation frequency.

4. The method for combined bioinformatics analysis of tumor-specific gene mutation and methylation according to claim 3, characterized in that: When there is no difference between the corresponding mutation sites in the mutation frequency dataset of cancer samples and the mutation frequency dataset of healthy samples, the reference threshold determination process of the mutation site is as follows: If the maximum mutation frequency in the healthy sample mutation frequency data set of the mutation site is greater than or equal to 0.40, the reference threshold of the mutation site is set to 0.01; If the maximum mutation frequency in the healthy sample mutation frequency data set of the mutation site is less than 0.01 and 1.30 times the 0.99 quantile is also less than 0.01, the reference threshold of the mutation site is set to 0.01; If the maximum mutation frequency in the mutation frequency data set of healthy samples of the mutation site is greater than 0.01 and less than 0.40, and 1.30 times the 0.99 quantile is also greater than 0.01, then if the maximum value is greater than 1.30 times the 0.99 quantile, the reference threshold of the mutation site is set to the maximum mutation frequency of healthy samples in the training set; if the maximum value is less than or equal to 1.30 times the 0.99 quantile, the reference threshold of the mutation site is set to 1.30 times the 0.99 quantile of the mutation frequency of healthy samples in the training set; If the maximum mutation frequency in the healthy sample mutation frequency data set for the mutation site is greater than 0.01 and less than 0.40, and 1.30 times the 0.99 quantile is less than or equal to 0.01, the reference threshold of the mutation site is set to the maximum mutation frequency of the healthy sample in the training set; If the maximum mutation frequency in the healthy sample mutation frequency data set of the mutation site is less than or equal to 0.01 and 1.30 times the 0.99 quantile is greater than 0.01, the reference threshold of the mutation site is set to 1.30 times the 0.99 quantile of the mutation frequency of healthy samples in the training set.

5. The method for combined bioinformatics analysis of tumor-specific gene mutation and methylation according to claim 1, characterized in that: The Ω value indicates the closeness between the mutation frequency of the mutation site in the sample to be tested and the cancer sample. The overall Ω value of the sample to be tested is the sum of the Ω values ​​of each mutation site in the sample to be tested. The process of calculating the Ω value of each mutation site in the sample to be tested is as follows: If the mutation frequency of the mutation site in the sample to be tested is greater than or equal to 0.40, set Ω = 1; If the mutation frequency of the mutation site in the sample to be tested is less than or equal to the reference threshold, and if the reference threshold is less than 0.01, let Ω = 99.999*AF; If the reference threshold is greater than or equal to 0.01, let Ω = 50*AF; If the mutation frequency of the mutation site of the sample to be tested is greater than the reference threshold, if the reference threshold is less than 0.01 and the mutation frequency of the sample to be tested is greater than 0.005, then Ω=99.999*AF+1; if the reference threshold is less than 0.01 and the mutation frequency of the sample to be tested is less than or equal to 0.005, then Ω=99.999*AF; if the reference threshold is greater than or equal to 0.01 and the mutation frequency of the sample to be tested is greater than 0.005, then Ω=50*AF+1; if the reference threshold is greater than or equal to 0.01 and the mutation frequency of the sample to be tested is less than or equal to 0.005, then Ω=50*AF; Among them, the AF value is used to indicate the frequency of each allele at a mutation site; the AF value is a value between 0 and 1, where 0 indicates that the frequency of the allele among all alleles is 0, and 1 indicates that the frequency of the allele among all alleles is 1.

6. The method for combined bioinformatics analysis of tumor-specific gene mutation and methylation according to claim 1, characterized in that: The first condition is whether the proportion of methylation data with a Zscore value greater than 3 is greater than 7%, or whether the maximum Zscore value is greater than 10, or whether the average Zscore value is greater than 0.4; The second condition is whether the proportion of Zscore values ​​greater than 1 in the methylation data is less than 3%, or whether the maximum Zscore value is less than 1.80, or whether the average Zscore value is less than -1.

05.

7. The method for combined bioinformatics analysis of tumor-specific gene mutation and methylation according to claim 1, characterized in that: In the random forest auxiliary model constructed based on methylation data, if the probability that the sample to be tested is a healthy sample is not less than 65%, the type of the sample to be tested is determined to be a healthy type; if the probability that the sample to be tested is a healthy sample is less than 65%, a fifth judgment is made based on the neural network model constructed based on methylation data.

8. The method for combined bioinformatics analysis of tumor-specific gene mutation and methylation according to claim 1, characterized in that: In the neural network model constructed based on methylation data, if the probability value of the sample to be tested is a healthy sample is greater than 50%, the type of the sample to be tested is determined to be a healthy type; if the probability value of the sample to be tested is not greater than 50%, the type of the sample to be tested is determined to be a cancer type.

9. The method for combined bioinformatics analysis of tumor-specific gene mutation and methylation according to claim 1, characterized in that: The process of calculating the Zscore value of the methylation data of the sample to be tested is as follows: Set a training set, which includes cancer samples and healthy samples. The cancer samples include samples of different types of cancer patients; Calculate the mean μ and standard deviation σ of each methylation site in the healthy samples in the training set; Calculate the Zscore value of the test sample compared to the training set. The calculation formula is: Wherein, X represents the methylation level value of the corresponding methylation site of the sample to be tested.

10. A tumor-specific gene mutation and methylation combined bioinformatics analysis system, characterized in that: It includes a data processing module, a mutation site detection module, a methylation detection module and a cancer type detection module, wherein the methylation detection module includes a Zscore value detection model, a random forest auxiliary model and a neural network model; The data processing module is used to perform upstream analysis on the sample to be tested to obtain data on each mutation site of the sample and methylation data of the sample; The mutation site detection module is used to calculate the Ω value of each mutation site in the sample to be tested based on the reference threshold of the mutation site detection model and output whether the sample to be tested is a cancer type according to the overall Ω value; The methylation detection module is used to judge the methylation data of the sample to be tested through the Zscore value detection model, the random forest auxiliary model, and the neural network model in sequence, and output whether the sample to be tested is a cancer type; The cancer type detection module is used to determine which cancer type the sample to be tested belongs to based on methylation data when the mutation site detection module or the methylation detection module outputs that the sample to be tested is a cancer type.

Citation Information

Patent Citations

  • Method of mutation detection in a liquid biopsy

    EP4130293A1

  • Anomalous fragment detection and classification

    US20190287652A1

  • Ultra-sensitive detection of circulating tumor DNA through genome-wide integration

    US20210043275A1

  • Tumor biomarker, and cancer risk information generation method and apparatus

    WO2024046207A1

  • Bladder cancer methylation site marker and use thereof

    WO2024222786A1