Method and system for generating and predicting cancer risk level prediction model
By combining methylation site data with gold standard multi-model training, a cancer risk level prediction model was generated, which solved the problems of PSA testing errors and overdiagnosis, and achieved the accuracy of cancer diagnosis and the rational allocation of resources.
Patent Information
- Application Number
- PCT/CN2024/085579
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-02
- Publication Date
- 2025-10-09
AI Technical Summary
Existing cancer screening methods such as PSA testing have errors and cannot distinguish between benign and malignant tumors, leading to overdiagnosis and overtreatment. They have low accuracy rates, and methylation testing is not widely used in clinical practice.
By combining cancer-related methylation site data with the gold standard, a cancer risk level prediction model is generated through iterative training and prediction of multiple machine learning models, which is used to accurately diagnose and predict cancer risk levels.
It improves the accuracy of cancer diagnosis, helps doctors and patients better understand the diagnosis results, rationally allocate medical resources, reduce unnecessary examinations and treatments, and improve the accuracy of disease risk prediction.
Smart Images

Figure CN2024085579_09102025_PF_FP_ABST
Abstract
Description
Method and system for generating and predicting cancer risk level prediction model Technical Field
[0001] This specification relates to the field of biotechnology, and in particular to a method and system for generating and predicting a cancer risk level prediction model. Background Art
[0002] Cancer is the leading cause of death worldwide, and the number of cancer deaths each year shows a clear upward trend. Taking prostate cancer as an example, it has become the fifth most common cancer among Chinese men. Early prostate cancer has no clinical symptoms, so most patients are already in the middle or late stages of the disease when diagnosed, and there is little hope of cure. Currently, cancer screening is mainly carried out through some diagnostic markers. For example, prostate cancer is detected by serum prostate specific antigen (PSA) testing. However, PSA testing has the following problems: (1) Errors in measurement results: The level of PSA can be affected by many factors, such as prostatitis and prostatic hyperplasia, so its measurement results may be erroneous. (2) Unable to distinguish between benign and malignant: PSA testing cannot distinguish between benign hyperplasia and malignant tumors, so even if the PSA level is elevated, it cannot be determined whether prostate cancer has occurred. (3) Overdiagnosis and overtreatment: Due to errors in PSA testing, overdiagnosis and overtreatment may occur, which brings unnecessary physical, psychological and economic burdens to patients. At the same time, it increases the difficulty of doctors' work and reduces their work efficiency. Unilateral PSA testing cannot diagnose and predict prostate cancer well, and the accuracy rate is not high.
[0003] In cancer, altered methylation patterns are often associated with tumor development and progression.
[0004] Therefore, it is desirable to provide a method and system for diagnosing and predicting cancer by utilizing methylation site data associated with cancer and combining it with the gold standard for cancer diagnosis.
[0005] Summary of the Invention
[0006] One or more embodiments of the present specification provide a method for generating a cancer risk level prediction model, wherein the prediction model includes a first prediction model, a second prediction model, a third prediction model, a fourth prediction model, and a fifth prediction model, and the method includes: step S1, using a training set and a label to train an initial first prediction model, an initial second prediction model, an initial third prediction model, and an initial fourth prediction model respectively to obtain trained first prediction model, second prediction model, third prediction model, and fourth prediction model, wherein the training set includes methylation data of cancer-related methylation sites of each sample in the training sample, and the label is that the sample has the cancer risk level; step S2, using the trained first prediction model, second prediction model, third prediction model, and fourth prediction model respectively to predict the samples in the prediction set to obtain respective prediction results of the four prediction models, wherein the prediction results of the third prediction model and the fourth prediction model are input into the fifth prediction model to obtain the prediction result of the fifth prediction model; and step S3, judging whether each of the prediction results meets a preset condition, and if so, generating each prediction model; if not, returning to step S1 and further iterating until each prediction result meets the preset condition.
[0007] One or more embodiments of the present specification provide a system for determining a prediction model capable of predicting the cancer risk level of a subject, comprising: a training module for performing the following operations: step S1, using a training set and a label to respectively train an initial first prediction model, an initial second prediction model, an initial third prediction model, and an initial fourth prediction model to obtain a trained first prediction model, a second prediction model, a third prediction model, and a fourth prediction model, wherein the training set includes methylation data of cancer-related methylation sites of each sample in the training sample, and the label is that the sample has the cancer risk level; step S2, using the trained first prediction model, the second prediction model, the third prediction model, and the fourth prediction model to respectively predict the samples in the prediction set to obtain respective prediction results of the four prediction models, wherein the prediction results of the third prediction model and the fourth prediction model are input into the fifth prediction model to obtain the prediction result of the fifth prediction model; and step S3, judging whether each prediction result meets a preset condition, and if so, generating each prediction model; if not, returning to step S1 for further iteration until each prediction result meets the preset condition.
[0008] One or more embodiments of the present specification provide a device for determining a prediction model capable of predicting a subject's risk level of cancer, comprising at least one processor and at least one memory; the at least one memory is used to store computer instructions; and the at least one processor is used to execute at least part of the computer instructions to implement the above-mentioned method of generating a cancer risk level prediction model.
[0009] One or more embodiments of this specification provide a computer-readable storage medium that stores computer instructions. When a computer reads the computer instructions in the storage medium, the computer executes the above-mentioned method for generating a cancer risk level prediction model.
[0010] One or more embodiments of the present specification provide a method for predicting the risk level of a subject suffering from cancer, comprising: obtaining methylation data of methylation sites associated with the cancer; inputting the methylation data of the subject into a prediction model to predict the risk level of the subject suffering from cancer, wherein the prediction model is a first prediction model, a second prediction model, a third prediction model, a fourth prediction model, or a fifth prediction model, wherein the prediction model is generated by the following method: step S1, using a training set and a label to respectively train an initial first prediction model, an initial second prediction model, an initial third prediction model, and an initial fourth prediction model, to obtain the trained first prediction model, the second prediction model, the third prediction model, and the fourth prediction model. The model comprises a training set including methylation data of cancer-related methylation sites of each sample in the training samples, and a label indicating that the sample has the cancer risk level; step S2, respectively using the trained first prediction model, second prediction model, third prediction model and fourth prediction model to predict the samples in the prediction set, and obtaining respective prediction results of the four prediction models, wherein the prediction results of the third prediction model and the fourth prediction model are input into the fifth prediction model to obtain the prediction result of the fifth prediction model; and step S3, determining whether each prediction result meets the preset conditions, and if so, generating each prediction model; if not, returning to step S1 for further iteration until each prediction result meets the preset conditions.
[0011] One or more embodiments of the present specification provide a system for predicting the risk level of a subject suffering from cancer, comprising: an acquisition module for acquiring methylation data of methylation sites associated with the cancer; a prediction module for inputting the methylation data of the subject into a prediction model to predict the risk level of the subject suffering from cancer, wherein the prediction model is generated by: step S1, using a training set and labels to respectively train an initial first prediction model, an initial second prediction model, an initial third prediction model, and an initial fourth prediction model to obtain trained first prediction models, second prediction models, third prediction models, and fourth prediction models, wherein the training set includes training samples The methylation data of the cancer-related methylation sites of each sample in this method are labeled with the cancer risk level of the sample; step S2, using the trained first prediction model, second prediction model, third prediction model and fourth prediction model to predict the samples in the prediction set, respectively, to obtain the prediction results of the four prediction models, wherein the prediction results of the third prediction model and the fourth prediction model are input into the fifth prediction model to obtain the prediction result of the fifth prediction model; and step S3, judging whether each prediction result meets the preset conditions. If so, the training is completed; if not, returning to step S1 for further iteration until each prediction result meets the preset conditions.
[0012] One or more embodiments of the present specification provide a device for predicting a subject's risk level of developing cancer, comprising at least one processor and at least one memory; the at least one memory is used to store computer instructions; and the at least one processor is used to execute at least part of the computer instructions to implement the above-mentioned method for predicting a subject's risk level of developing cancer.
[0013] One or more embodiments of the present specification provide a computer-readable storage medium storing computer instructions. When a computer reads the computer instructions in the storage medium, the computer executes the above-mentioned method for predicting the risk level of a subject suffering from cancer. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] This specification will be further described in the form of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting, and in these embodiments, like numbers represent like structures, wherein:
[0015] FIG1 is a flowchart of prostate cancer analysis based on combined PSA and methylation-specific profiles according to some embodiments of the present specification;
[0016] FIG2 is a diagram of a model hyperparameter solution and analysis process according to some embodiments of this specification;
[0017] 3A and 3B are six-category cluster diagrams based on combined PSA and methylation-specific profiles according to some embodiments of the present specification;
[0018] FIG4 is a methylation heat map of methylation sites (14) screened according to the training set described in some embodiments of this specification;
[0019] FIG5A and FIG5B are methylation heat maps of methylation sites (14) screened according to the training set described in some embodiments of this specification;
[0020] FIG6A and FIG6B are methylation heat maps of methylation sites (28) screened according to the training set described in some embodiments of this specification;
[0021] FIG7 is a performance diagram of a confusion matrix for predicting six categories of a sample set using three evaluators at fPSA and tPSA, respectively, according to some embodiments of this specification;
[0022] FIG8 is a Roc curve diagram plotted based on the likelihood probability risk levels of high (H), medium (M) and low (L) risk levels given by the KNN algorithm based on the extracted Bayesian evaluator according to some embodiments of this specification;
[0023] FIG9 is a confusion matrix diagram drawn based on Bayesian risk level prediction results according to some embodiments of this specification;
[0024] FIG10 is a confusion matrix diagram of prostate cancer and non-prostate cancer prediction using three classifiers according to some embodiments of this specification;
[0025] FIG11 is a Venn diagram of prediction consistency based on three evaluators for prediction results of fPSA vs. tPSA of prediction set samples according to some embodiments of this specification;
[0026] FIG12 is a schematic diagram of an application scenario of a cancer risk level prediction system according to some embodiments of this specification;
[0027] FIG13 is a module diagram of a cancer risk level prediction system according to some embodiments of this specification;
[0028] FIG14 is a flowchart of a method for determining a cancer risk level prediction model according to some embodiments of this specification;
[0029] FIG15 is a flowchart of a method for predicting cancer risk level according to some embodiments of the present specification. DETAILED DESCRIPTION
[0030] To more clearly illustrate the technical solutions of the embodiments of this specification, the following briefly describes the drawings required for describing the embodiments. Obviously, the drawings described below are merely examples or embodiments of this specification. Those skilled in the art can apply this specification to other similar scenarios based on these drawings without inventive effort. Unless otherwise apparent from the context or otherwise noted, the same reference numerals in the figures represent the same structure or operation.
[0031] It should be understood that the terms "system," "device," "unit," and / or "module" used herein are a method for distinguishing different components, elements, parts, portions, or assemblies at different levels. However, if other terms can achieve the same purpose, the terms may be replaced by other expressions.
[0032] As used in this specification and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" do not refer to the singular but also include the plural. Generally speaking, the terms "comprises" and "include" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include other steps or elements.
[0033] Flowcharts are used throughout this specification to illustrate the operations performed by systems according to embodiments of this specification. It should be understood that preceding or following operations do not necessarily need to be performed in exact order. Instead, the steps may be processed in reverse order or simultaneously. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.
[0034] Methylation is an epigenetic modification that involves the addition and removal of methyl groups on DNA molecules. In cancer, changes in methylation patterns are often associated with the development and progression of tumors. However, it should be noted that methylation testing is currently mainly used in research and laboratory settings and is not widely used for early cancer screening in clinical practice. This is because methylation testing technology still has some challenges, such as standardization and reproducibility, and the interpretation of methylation test results also requires further research and verification. Cancer markers can be used to screen for cancer, but there are problems such as errors and overdiagnosis.
[0035] Therefore, the present invention has developed a method for screening and diagnosing cancer by combining markers and cancer-specific methylation data. Specifically, cancer is classified as a label using markers and / or clinical diagnostic results, and cancer-specific methylation data is used as training data to train a machine learning model, thereby achieving accurate diagnosis of cancer. Cancer can be classified into multiple categories, for example, multiple classifications (e.g., high, medium, and low) are performed on patients who may have cancer, and multiple classifications (e.g., high, medium, and low) are also performed on patients who may not have cancer. This can help doctors and patients better understand the diagnostic results and take appropriate preventive or therapeutic measures, thereby improving the accuracy of prediction of the risk of disease; it is helpful to rationally allocate medical resources, so that high-risk groups can get more attention and resource support, while low-risk groups can avoid unnecessary examinations and treatments, saving resources. This diagnostic scheme has huge clinical promotion and application value.
[0036] FIG12 is a schematic diagram of an application scenario of a cancer risk level prediction system according to some embodiments of this specification. As shown in FIG12 , an application scenario 100 of a cancer risk level prediction system may include a processing device 110 , a storage device 120 , a network 130 , a user terminal 140 , and a detection device 160 .
[0037] The processing device 110 can be used to process relevant data and / or information of the cancer risk level prediction system. In some embodiments, the processing device 110 can obtain data and / or information from the storage device 120 or other components of the application scenario 100 of the cancer risk level prediction system (for example, the user terminal 140, the detection device 160), and execute program instructions based on this information and / or data to perform one or more functions described in this specification. For example, the processing device 110 can obtain training set data from the storage device 120 and train a prediction model based on the training set data. For another example, the processing device 110 can obtain the original sequencing data of the clinical sample of the subject measured by the detection device 160, obtain methylation site data after data processing of the original sequencing data, and call the prediction model stored in the storage device 120 to process the methylation site data to predict the cancer risk level of the subject. In some embodiments, the processing device 110 can be a server or a central processing unit.
[0038] The storage device 120 can be used to store data, instructions and / or any other information. In some embodiments, the storage device 120 can store data and / or information obtained from the processing device 110 or other components of the application scenario 100 of the cancer risk level prediction system (e.g., the user terminal 140, the detection device 160). For example, the storage device 120 can store various prediction models for use by the processing device 110. For another example, the storage device 120 can obtain and store the original sequencing data of the subject from the detection device 160. For another example, the storage device 120 can receive and store information uploaded by the user terminal 140, such as the subject's clinical information.
[0039] Network 130 may comprise any suitable network capable of facilitating information and / or data exchange within application scenario 100 of the cancer risk level prediction system. In some embodiments, processing device 110 and other components of application scenario 100 of the cancer risk level prediction system (e.g., storage device 120, user terminal 140, and detection device 160) may exchange information via network 130. For example, processing device 110 may receive data from storage device 120 via network 130. For another example, raw sequencing data of a subject measured by detection device 160 may be transmitted to processing device 110 via the network. In some embodiments, network 130 may comprise one or more of a wired network or a wireless network. For example, network 130 may comprise a cable network, a fiber optic network, or the like. In some embodiments, network 130 may comprise various topologies, such as point-to-point, shared, or centralized, or a combination of multiple topologies. In some embodiments, network 130 may comprise one or more network access points. For example, via access points such as base stations and / or one or more network switching points, one or more components of application scenario 100 may connect to network 130 to exchange data and / or information.
[0040] User terminal 140 can be used to implement the services provided to users by application scenario 100 of the cancer risk level prediction system. For example, a user can send a subject's clinical test information to processing device 110 via user terminal 140. For another example, a user can send raw sequencing data of a subject's DNA sample to processing device 110 via user terminal 140. For example, user terminal 140 can display the subject's cancer risk level predicted by the prediction model. In some embodiments, user terminal 140 can include one or any combination of a smartphone 140-1, a tablet computer 140-2, a laptop computer 140-3, or other devices with input and / or output functions.
[0041] The detection device 160 is used to detect raw sequencing data from a clinical sample 150 from a subject. Exemplary clinical samples 150 may include, but are not limited to, blood, urine, sweat, and the like. As examples, the detection device 160 may include devices that implement one or more of the following methods for determining methylation data: targeted methylation sequencing, WGBS, RRBS, oxBS-seq, MethylCap-seq, MBD-seq, MeDIP-seq, HPLC, MSRF, MASP, methylation chip methods, pyrosequencing, dPCR, and MS-PCR. In some embodiments, the detection device 160 may include a DNA sequencer, a methylation chip, and the like.
[0042] It should be noted that the application scenarios are provided for illustrative purposes only and are not intended to limit the scope of this specification. A person skilled in the art can make various modifications or variations based on the description of this specification. For example, the application scenarios may also include a database. For another example, the application scenarios may be implemented on other devices to achieve similar or different functions. However, such changes and modifications do not deviate from the scope of this specification.
[0043] FIG13 is a module diagram 200 of a cancer risk level prediction system according to some embodiments of the present specification.
[0044] In some embodiments, the module diagram 200 of the cancer risk level prediction system may include an acquisition module 210, a training module 220, and a prediction module 230. In some embodiments, these modules may be implemented by the processing device 110 in FIG. 1 .
[0045] In some embodiments, the acquisition module 210 is configured to acquire training sample data (e.g., first sample data, second sample data, third sample data, fourth sample data) and label data (e.g., first label, second label, third label, fourth label) for the cancer risk level prediction system. In some embodiments, the training sample data may include methylation data of cancer-related methylation sites for each sample in the training sample, and the methylation data may include the methylation level of the methylation site, the average methylation level, the sequencing depth of each methylation site, the average sequencing depth, etc.
[0046] In some embodiments, the training module 220 is configured to obtain the trained first, second, third and fourth prediction models based on the sample data and label data through the initial first, second, third and fourth prediction models, and obtain the prediction results based on the trained prediction models, and judge whether each prediction result meets the preset conditions, thereby generating a prediction model or further returning to iteration.
[0047] In some embodiments, the training module 220 may be configured to obtain a trained first prediction model based on the first training sample and the first label using the initial first prediction model, and obtain a first prediction result based on the trained first prediction model; obtain a trained second prediction model based on the second training sample and the second label using the initial second prediction model, and obtain a second prediction result based on the trained second prediction model; obtain a trained third prediction model based on the third training sample and the third label using the initial third prediction model, and obtain a third prediction result based on the trained third prediction model; and obtain a trained fourth prediction model based on the fourth training sample and the fourth label using the initial fourth prediction model, and obtain a fourth prediction result based on the trained fourth prediction model. In some embodiments, the training module 220 may be further configured to obtain a fifth prediction result based on the third prediction result and the fourth prediction result using the fifth prediction model.
[0048] In some embodiments, the training module 220 may be further configured to generate respective prediction models in response to the first prediction result, the second prediction result, the third prediction result, the fourth prediction result, and the fifth prediction result all satisfying preset conditions.
[0049] In some embodiments, the prediction module 230 may be configured to predict the cancer risk level of the subject.
[0050] For more details about the acquisition module 210, the training module 220 and the prediction module 230, please refer to Figure 14 and its related description.
[0051] It should be understood that the system and its modules shown in FIG13 can be implemented in various ways. It should be noted that the above description of the cancer risk level prediction system and its modules is for ease of description only and does not limit this specification to the scope of the illustrated embodiments. It is understood that, after understanding the principles of the system, those skilled in the art may arbitrarily combine the modules or form subsystems connected to other modules without departing from these principles. In some embodiments, the acquisition module 210, training module 220, and prediction module 230 disclosed in FIG13 may be different modules in a single system, or a single module may implement the functions of two or more of the aforementioned modules. In some embodiments, one or more modules may be omitted. For example, the training module 220 may be omitted, and model training may be completed in another system (e.g., a processing device) and then directly invoked by the cancer risk level prediction system. For example, the modules may share a storage module, or each module may have its own storage module. Such variations are within the scope of this specification.
[0052] Figure 14 is a flow chart of a method for generating a cancer risk level prediction model according to some embodiments of this specification. In some embodiments, process 1400 may be executed by processing device 110. As shown in Figure 14, process 1400 may include steps S1, S2, S3, and S4.
[0053] Step S1: Use the training set and labels to train the initial first prediction model, the initial second prediction model, the initial third prediction model and the initial fourth prediction model respectively to obtain the trained first prediction model, the second prediction model, the third prediction model and the fourth prediction model.
[0054] The cancer risk level prediction model (or simply the prediction model) includes a first, second, third, fourth, and fifth prediction models. As shown in the dashed box S1 in Figure 14, to generate the prediction model, it is necessary to first train the initial first, second, third, and fourth prediction models to obtain the trained first, second, third, and fourth prediction models.
[0055] The first prediction model, the second prediction model, the third prediction model, and the fourth prediction model can each be a machine learning model for predicting cancer risk levels. Exemplary machine learning models include, but are not limited to, K-nearest neighbor (KNN), decision tree models, random forest models, support vector machine models (SVM), naive Bayesian models, logistic regression models, artificial neural network models, gradient boosting models, multilayer perceptron models, elastic network models, and the like, or combinations thereof. In some embodiments, the first and third prediction models can be the same model, and the second and fourth prediction models can be the same model. It should be noted that the terms "first," "second," "third," "fourth," and "fifth" in this specification are merely used to distinguish between different prediction models and should not be construed as limiting. As long as the prediction model generation method described in this specification can be implemented, the "first," "second," "third," "fourth," and "fifth" prediction models can be arbitrarily combined, arranged, or swapped in order. For example, the third prediction model and the fourth prediction model can be swapped in order. For example, the first and fourth prediction models can be the same model, and the second and third prediction models can be the same model. In some embodiments, the first and third prediction models can be KNN models, and the second and fourth prediction models can be SVM models. In some embodiments, the first and third prediction models may be SVM models, and the second and fourth prediction models may be KNN models. In some embodiments, the first and third prediction models may be decision tree models, and the second and fourth prediction models may be multilayer perceptron models.
[0056] In some embodiments, cancer may include any cancer that can be predicted and classified using a marker, including but not limited to prostate cancer, colorectal cancer, gastric cancer, liver cancer, testicular cancer, ovarian cancer, pancreatic cancer, gallbladder cancer, breast cancer, head and neck squamous cell carcinoma, bladder cancer, etc. By way of example only, the marker for prostate cancer is prostate-specific antigen (PSA); the marker for liver cancer and testicular cancer is alpha-fetoprotein; the marker for colon cancer and gastric cancer is carcinoembryonic antigen; the marker for ovarian cancer is cancer antigen 125; the marker for breast cancer is cancer antigen 15-3; the markers for pancreatic cancer and gallbladder cancer are carbohydrate antigen 19-9, carbohydrate antigen 242, and cancer antigen 50; the marker for head and neck squamous cell carcinoma is squamous cell carcinoma antigen; and the marker for bladder cancer is nuclear matrix protein-22.
[0057] The cancer risk level can be assessed and classified based on the likelihood that the subject will develop cancer. In this specification, the cancer risk level has different levels according to different judgment criteria. In some embodiments, the cancer risk level can be determined based on clinical classification. Specifically, the cancer risk level can be determined as cancer and non-cancer based on the pathological diagnosis results of the hospital (for example, CT scan, ultrasound, etc.), that is, the cancer risk level in this case is two-category. In some embodiments, the cancer risk level can be determined based on the numerical value of the cancer marker. Specifically, the cancer risk level of the subject can be determined based on the comparison of the numerical value of the marker with the reference numerical value. For different markers, the cancer risk level can be a high probability, a medium probability and a low probability, or the cancer risk level can be a high probability and a low probability, that is, the cancer risk level in this case is three-category or two-category.
[0058] For prostate cancer, the marker PSA exists in the blood in two forms: bound PSA (cPSA) and free PSA (fPSA). The PSA in the blood is the sum of free PSA and compound PSA, also known as total PSA, expressed as tPSA. tPSA or fPSA / tPSA can be used to confirm the risk level of prostate cancer. When the value of the prostate cancer marker is tPSA, if tPSA is less than a first concentration value (e.g., 4 ng / mL), it indicates that the cancer risk level is low, that is, the risk of having prostate cancer is low; if tPSA is between the first concentration value (e.g., 4 ng / mL) and the second concentration value (e.g., 10 ng / mL), the cancer risk level is medium, that is, the risk of having prostate cancer is medium; if tPSA is greater than the second concentration value (e.g., 10 ng / mL), the sample has a high risk of having the cancer, that is, the risk of having prostate cancer is high. When the value of the prostate cancer marker is fPSA / tPSA, if fPSA / tPSA is greater than a first value (e.g., 0.25), the cancer risk level is low, that is, the risk of having prostate cancer is low; if fPSA / tPSA is between the first value (e.g., 0.25) and the second value (e.g., 0.16), the cancer risk level is medium, that is, the risk of having prostate cancer is medium; if fPSA / tPSA is less than the second value (e.g., 0.16), the cancer risk level is high, that is, the risk of having prostate cancer is high. The prostate cancer risk level in this case is divided into three categories.
[0059] For colon cancer, the normal serum CEA marker value is <5 μg / L. If the CEA concentration is greater than or equal to 5 μg / L, the cancer risk level is high, indicating a higher risk of colon and stomach cancer. If the CEA concentration is less than 5 μg / L, the cancer risk level is low, indicating a lower risk of colon and stomach cancer. In this case, the colon cancer risk level is binary.
[0060] In some embodiments, the cancer risk level can be determined based on the values of the cancer markers in the sample and the clinical classification. Specifically, subjects clinically classified as cancer and non-cancer can be further classified based on the values of the markers, thereby obtaining cancer risk levels of high cancer risk, moderate cancer risk, low cancer risk, high non-cancer risk, moderate non-cancer risk, and low non-cancer risk, i.e., in this case, the cancer risk level is divided into six categories, or the cancer risk level is divided into high cancer risk, low cancer risk, high non-cancer risk, and low non-cancer risk, i.e., in this case, the cancer risk level is divided into four categories.
[0061] For prostate cancer, if a subject is clinically diagnosed with prostate cancer and the fPSA / tPSA value is 0.177 (between the first value (e.g., 0.25) and the second value (e.g., 0.16), the cancer risk level is medium), then the cancer risk level determined based on the numerical value of the cancer marker in the sample and the clinical classification is medium cancer risk; if a subject is clinically diagnosed with prostate cancer and the tPSA concentration is 7.371 ng / mL (between the first concentration value (e.g., 4 ng / mL) and the second concentration value (e.g., 10 ng / mL), the cancer risk level is medium), then the cancer risk level determined based on the numerical value of the cancer marker in the sample and the clinical classification is medium cancer risk. If a subject is clinically diagnosed as not having prostate cancer and the fPSA / tPSA value is 0.379 (greater than the first value (for example, 0.25), the cancer risk level is low probability), then the cancer risk level determined based on the value of the cancer marker in the sample and the clinical classification is low probability of non-cancer; if a subject is clinically diagnosed as not having prostate cancer and the tPSA concentration is 0.496 ng / mL (less than the first concentration value (for example, 4 ng / mL), the cancer risk level is low probability), then the cancer risk level determined based on the value of the cancer marker in the sample and the clinical classification is low probability of non-cancer.
[0062] For colon cancer, if a subject is clinically diagnosed with colon cancer and the concentration of the marker CEA is greater than or equal to 5 μg / L, then the cancer risk level determined based on the value of the cancer marker in the sample and the clinical classification is a high probability of cancer; if a subject is clinically diagnosed with colon cancer and the concentration of the marker CEA is less than 5 μg / L, then the cancer risk level determined based on the value of the cancer marker in the sample and the clinical classification is a low probability of cancer.
[0063] The training set for training each prediction model includes the methylation data of the methylation sites of each sample in the training sample, and the label is that the sample has the cancer risk level. In some embodiments, the sample may include subjects with cancer and subjects without cancer. Subjects without cancer include healthy subjects, subjects with diseases that interfere with cancer prediction, etc. As an example only, for prostate cancer, subjects without cancer may include healthy subjects, subjects with other prostate problems (such as benign hyperplasia, prostate cysts) or subjects with other cancers (e.g., bladder cancer).
[0064] The methylation sites in the training set used to train different prediction models are obtained by screening from cancer-related methylation sites using a classification algorithm. Cancer-related methylation sites refer to sites with significant methylation differences between cancer tissues and normal tissues. For example, these methylation sites are expressed at higher levels in cancer tissues, but expressed at lower levels or not expressed in normal tissues. Cancer-related methylation sites and their methylation data can be obtained through high-throughput sequencing, for example, targeted methylation sequencing. Methylation data may include the methylation rate, average methylation level, sequencing depth, average sequencing depth, etc. of each methylation site. There are many cancer-related methylation sites, but not all of these methylation sites are suitable for model training. It is necessary to screen out methylation sites that are suitable for each model. In some embodiments, the classification algorithm is a machine learning model that can screen out suitable model training from a variety of methylation sites. Exemplary classification algorithms include random forest algorithms, adaptive boosting algorithms, gradient boosting algorithms, etc. In some embodiments, all training set samples are first divided into two categories, three categories, two categories, six categories, or four categories according to the cancer risk level classification method described above; then, a series of cancer-related methylation sites of each sample can be input into the classification algorithm according to each category. For each risk level classification, the classification algorithm outputs a contribution value for each methylation site, and these contribution values are sorted from large to small. The methylation sites with the highest ranking (for example, top 10, top 15, top 20, top 30, etc.) are confirmed as the methylation sites in the group of risk level classifications.
[0065] In some embodiments, the first and second prediction models use the same labels and the same training set data. The labels of the first and second prediction models are divided into six categories or four categories for the risk level of the sample suffering from the cancer, namely, high probability of cancer, medium probability of cancer, low probability of cancer, high probability of non-cancer, medium probability of non-cancer and low probability of non-cancer, or high probability of cancer, low probability of cancer, high probability of non-cancer and low probability of non-cancer. As described above, the labels used by the first and second prediction models are determined based on the numerical value and clinical classification of the cancer marker in the sample.
[0066] When using tPSA to predict prostate cancer, the methylation sites used to train the first prediction model and the second prediction model include at least: IGFBP3_83, IGFBP3_16, IGFBP3_121, IGFBP3_48, IGFBP3_36, FEZF2_21, IGFBP3_44, IGFBP3_114, SERPINB1_4, FEZF2_74, FEZF2_5, IGFBP3_6 7. SERPINB1_67, SERPINB1_97, ZNF154_6, SOX1-OT_5, FEZF2_23, APC_78, FHAD1_57, SERPINB1_51 , APC_98, IGFBP3_95, MIR663A_3, inTerGrnTiC_1_84, SERPINB1_36, POU4F2_77, APC_36 and APC_110.
[0067] When using fPSA / tPSA to predict prostate cancer, the methylation sites used to train the first prediction model and the second prediction model include at least: IGFBP3_36, IGFBP3_121, SERPINB1_4, IGFBP3_89, FEZF2_5, SOX1-OT_3, MIR663A_55, FHAD1_27, FEZF2_74, SOX1-OT_58, MIR663A_61, MIR663A_90 , APC_105, POU4F2_77, POU4F2_98, FHAD1_2, MIR663A_50, inTerGrnTiC_1_69, APC_36, FHAD1_57, inTe rGrnTiC_1_53, SOX1-OT_5, SOX1-OT_74, SERPINB1_91, IGFBP3_67, APC_117, IGFBP3_101 and ZNF154_116.
[0068] The third and fourth prediction models use different labels but have the same, different, or overlapping training data. The third prediction model labels the sample's risk level for the cancer as high, medium, or low. The third prediction model labels are determined based on the values of the cancer markers in the sample. The fourth prediction model labels the sample's risk level for the cancer as cancerous or non-cancerous, as determined based on clinical testing.
[0069] When using tPSA to predict prostate cancer, the methylation sites used to train the third prediction model include at least: SERPINB1_4, IGFBP3_36, IGFBP3_121, IGFBP3_114, SERPINB1_67, FEZF2_21, IGFBP3_16, FHAD1_2, IGFBP3_101, FEZF2_31, SOX1-OT_105, IGFBP3_83, FEZF2_74 and IGFBP3_44.
[0070] When predicting prostate cancer using fPSA / tPSA, the methylation sites used to train the third prediction model include at least: IGFBP3_114, MIR663A_39, IGFBP3_16, ZNF154_95, SOX1-OT_89, FHAD1_27, APC_20, IGFBP3_44, FEZF2_23, IGFBP3_48, SOX1-OT_74, SOX1-OT_3, IGFBP3_36 and SOX1-OT_5.
[0071] It should be noted that the above methylation sites are only examples, and these methylation sites may also include other methylation sites that can be used to implement model prediction.
[0072] The methylation sites used to train the fourth prediction model include at least: FEZF2_21, IGFBP3_121, IGFBP3_114, MIR663A_3, IGFBP3_36, FHAD1_2, ZNF154_116, SOX1-OT_58, IGFBP3_16, FEZF2_66, SERPINB1_4, SERPINB1_34, APC_78 and IGFBP3_101.
[0073] In some embodiments, a first training sample with a first label, a second training sample with a second label, a third training sample with a third label, and a fourth training sample with a fourth label are respectively input into an initial first prediction model, an initial second prediction model, an initial third prediction model, and an initial fourth prediction model for training, and a loss function is respectively constructed using the first label and the result of the initial first prediction model, the second label and the result of the initial second prediction model, the third label and the result of the initial third prediction model, and the fourth label and the result of the initial fourth prediction model. Based on the loss function, the parameters of the initial first, second, third, and fourth prediction models are iteratively updated to obtain the trained first, second, third, and fourth prediction models.
[0074] Step S2, respectively use the trained first prediction model, second prediction model, third prediction model and fourth prediction model to predict the samples in the prediction set, and obtain the respective prediction results of the four prediction models, wherein the prediction results of the third prediction model and the fourth prediction model are input into the fifth prediction model to obtain the prediction results of the fifth prediction model.
[0075] The methylation data in the prediction set are respectively input into the trained first prediction model, the second prediction model, the third prediction model, and the fourth prediction model, which respectively output prediction results. In some embodiments, the first prediction model and the second prediction model output six-category or four-category prediction results determined based on the markers and clinical classification, i.e., the first prediction result and the second prediction result; the third prediction model outputs a three-category or two-category prediction result determined by the markers, i.e., the third prediction result; the fourth prediction model outputs a two-category prediction result determined by the clinical classification, i.e., the fourth prediction result. In some embodiments, the prediction result can be represented by the prediction value output by the prediction model. By way of example only, for the two-category result, 0 represents cancer and 1 represents non-cancer; for the three-category result, 0 represents a high probability of cancer, 1 represents a moderate probability of cancer, and 2 represents a low probability of cancer; for the six-category result, 0 represents a high probability of cancer, 1 represents a moderate probability of cancer, 2 represents a low probability of cancer, 3 represents a high probability of non-cancer, 4 represents a moderate probability of non-cancer, and 5 represents a low probability of non-cancer.
[0076] The fifth prediction model is a machine learning model for performing six-category or four-category predictions on cancer. In some embodiments, the fifth prediction model is a Bayesian classifier, a decision tree, a K-nearest neighbor algorithm, etc. The prediction results of the third prediction model and the fourth prediction model are input into the fifth prediction model to obtain the prediction results of the fifth prediction model. That is to say, the prediction results of the three-category or two-category and the prediction results of the two-category are input into the fifth prediction model, and the fifth prediction model can directly output the prediction results of the six-category or four-category. Therefore, the prediction results of the first, second and fifth prediction models are consistent for the prediction classification of the same cancer, all of which are six-category or four-category.
[0077] Step S3, determine whether each of the prediction results meets the preset conditions. If so, generate the prediction models, that is, step S4; if not, return to step S1 for further iteration until each prediction result meets the preset conditions.
[0078] In some embodiments, the preset conditions may include the accuracy of the prediction results of the first prediction model, the second prediction model and the fifth prediction model reaching a first threshold respectively; the prediction results of the first prediction model and the second prediction model are converted into corresponding to the prediction results of the third prediction model, and the accuracy of the converted prediction results and the prediction results of the third prediction model reach a second threshold respectively; the prediction results of the first prediction model and the second prediction model are converted into corresponding to the prediction results of the fourth prediction model, and the accuracy of the converted prediction results and the prediction results of the fourth prediction model reach a third threshold respectively; and the consistency of the prediction results of the first prediction model, the second prediction model and the fifth prediction model reaches a fourth threshold.
[0079] The accuracy of the prediction results can be determined based on the output of the prediction model for a sample and the classified cancer risk level for that sample (for example, comparing the two to see if they are consistent). For example, if the output of the prediction model is completely consistent with the classified results, then the accuracy of the prediction results is 100%. The consistency of the prediction results can be confirmed based on the output of the first, second, and fifth prediction models. If the output of these three prediction models is the same for all samples, then the consistency is 100%.
[0080] Since the prediction results of the third prediction model are three-class or two-class, and the prediction results of the first and second prediction models are six-class or four-class, the prediction results of the first and second prediction models can be merged into the same classification as the prediction results of the third prediction model (i.e., three-class or two-class), and then compared to determine the accuracy of the prediction results. Similarly, since the prediction results of the fourth prediction model are two-class, and the prediction results of the first and second prediction models are six-class or four-class, the prediction results of the first and second prediction models can be merged into the same classification as the prediction results of the third prediction model (i.e., two-class), and then compared to determine the accuracy of the prediction results.
[0081] The first, second, third, and fourth thresholds may be system defaults or user-set. For example only, the first, second, third, and fourth thresholds may be any value greater than or equal to 0.85, 0.88, 0.9, 0.92, or 0.95. In some embodiments, the first threshold may be 0.90; the second threshold may be 0.85; the third threshold may be 0.9; and the fourth threshold may be 0.9.
[0082] In some embodiments, the preset condition includes: the number of iterations reaches a preset iteration value. The preset iteration value is a system default value or is set by the user. For example, the preset iteration value is 900, 950, 1000, 1050, 1100, etc.
[0083] Step S4: Generate each prediction model.
[0084] The generated prediction models can be stored in the storage device 120, and the processing device can call each prediction model from the storage device 120. These prediction models can all be used to predict the risk level of cancer, but the classification of the risk level is different. Specifically, the first, second, and fifth prediction models are used to predict the six-category or four-category results of cancer, the third prediction model is used to predict the three-category and two-category results of cancer, and the fourth prediction model is used to predict the two-category results of cancer. In some embodiments, the corresponding prediction model can be selected according to the user's desired classification. For example, if the user wants to obtain a six-category or four-category prediction result, any one or more of the first, second, and fifth prediction models can be selected.
[0085] FIG15 is a flowchart of a method for predicting cancer risk level according to some embodiments of this specification. In some embodiments, process 1500 may be executed by processing device 110. As shown in FIG15, process 1500 may include steps 1501 and 1503.
[0086] Step 1501: Acquire methylation data of methylation sites associated with the cancer.
[0087] Cancer-associated methylation sites refer to sites with significant methylation differences between cancer tissue and normal tissue. Methylation data can include the methylation rate, average methylation level, sequencing depth, average sequencing depth, etc. of each methylation site. In some embodiments, the methylation data of the methylation site can be obtained by the detection device 160.
[0088] Step 1503 , inputting the methylation data of the subject into a prediction model to predict the risk level of the subject suffering from cancer, wherein the prediction model is the first prediction model, the second prediction model, the third prediction model, the fourth prediction model or the fifth prediction model.
[0089] In some embodiments, the methylation data of the cancer-related methylation sites can be input into any one of the first prediction model, the second prediction model, the third prediction model, the fourth prediction model, or the fifth prediction model to output a prediction value, thereby confirming the subject's cancer risk level based on the prediction value. In some embodiments, data of methylation sites corresponding to the training set of each model can be selected from the cancer-related methylation sites and input into each prediction model. The prediction values output by these models can confirm the subject's cancer risk level; further comparing the cancer risk levels of the three prediction models, so that the predicted results are more accurate. In some embodiments, all of the methylation data can also be input into the first prediction model, the second prediction model, and the fifth prediction model.
[0090] Example
[0091] Example 1 Determination of methylation sites and methylation rates
[0092] (1) Urine sample collection
[0093] A total of 194 samples of morning urine were collected from 89 prostate cancer-positive patients and 105 prostate cancer-negative patients (including 3 renal pelvis cancer, 6 bladder cancer, 3 renal malignancies, and 3 prostate cysts to serve as interference samples for analytical accuracy). The samples were stored in 50 mL urine DNA storage tubes containing 7.5 mL of additives. After collection, the samples were centrifuged at 4000 rpm for 10 minutes, the supernatant discarded, and the pellet washed with 1× PBS.
[0094] (2) DNA extraction from urine precipitate
[0095] 1. Add 180 μL of Buffer GTL to the above pellet and resuspend the pellet; then add 20 μL of Proteinase K and vortex to mix.
[0096] 2. Incubate at 56°C for 1 hour until the sample is completely dissolved, then continue incubation at 90°C for 1 hour. Centrifuge briefly to collect the solution on the tube wall to the bottom of the tube.
[0097] 3. Add 200 μL of Buffer GL and vortex to mix thoroughly. Add 200 μL of anhydrous ethanol and vortex to mix thoroughly. Centrifuge briefly to collect the solution on the tube walls to the bottom.
[0098] 4. Add all the solution obtained in the previous step to the silicon matrix material membrane placed in the centrifuge tube.
[0099] 5. Add 500 μL of Buffer GW1 containing anhydrous ethanol to the silicon-based material membrane, centrifuge at 12,000 rpm for 1 min, discard the waste liquid in the collection tube, and place the silicon-based material membrane back into the collection tube.
[0100] 6. Add 500 μL of Buffer GW2 containing anhydrous ethanol to the silicon-based material membrane, centrifuge at 12,000 rpm for 1 min, discard the waste liquid in the collection tube, and place the silicon-based material membrane back into the collection tube.
[0101] 7. Centrifuge at 12,000 rpm for 2 minutes, discard the waste liquid in the collection tube, and place the silicon matrix material membrane at room temperature for several minutes to completely dry.
[0102] 8. Place the silica-based membrane in a new centrifuge tube, add 50-200 μL of Buffer GE, incubate at room temperature for 2-5 minutes, centrifuge at 12,000 rpm for 1 minute, collect the DNA solution, store the DNA at -20°C, and measure the DNA concentration using a Nano-300 microspectrophotometer and Qubit (the concentration should be no less than 1 ng / μL).
[0103] (3) Sulfite treatment of urine nucleic acid precipitation
[0104] 1. Prepare the DNA sample to be processed. The optimal DNA processing amount range is 20-1000 ng.
[0105] 2. Prepare Bisulfite Mix: Add 1.2 mL of MBuffer A-Conversion Solution to a tube of dry powder containing sodium bisulfite and shake until the dry powder is completely dissolved.
[0106] 3. Prepare the sulfite reaction system in a PCR tube: 50 μL urine precipitated DNA sample, 150 μL Bisulfite Mix, and 25 μL MBuffer B-protection solution.
[0107] 4. Sulfite conversion: After a brief centrifugation, place the PCR tube in a PCR instrument and operate: incubate at 85℃ for 50 minutes, then cool to room temperature and briefly centrifuge.
[0108] 5. DNA purification after sulfite treatment:
[0109] 5.1 Pour all the solution in the above PCR tube into a 1.5mL centrifuge tube;
[0110] 5.2 Add 285 μL MBuffer C-binding solution, 115 μL isopropanol, and 10 μL magnetic bead suspension to the centrifuge tube (please mix thoroughly before use) and shake for 10 minutes; after a brief centrifugation, place on a magnetic stand for adsorption for 2 minutes and discard the supernatant;
[0111] 5.3 Add 1000 μL MBuffer D-washing solution, incubate for 30 seconds without removing the magnetic stand, and discard the supernatant;
[0112] 5.4 Add 1000 μL MBuffer E-incubation solution, incubate at room temperature for 15 minutes, centrifuge briefly, place on a magnetic rack for 2 minutes, and discard the supernatant;
[0113] 5.5 Add 1000 μL MBuffer D-washing solution, incubate for 30 seconds without leaving the magnetic stand, and discard the supernatant; repeat this step once;
[0114] 5.6 After absorbing the excess detergent, place it on a clean bench and blow dry for 5 minutes.
[0115] 6. DNA purification and recovery after sulfite treatment:
[0116] 6.1 Add 50 μL of MBuffer F-elution buffer to the centrifuge tube and warm it at 56°C to improve the elution efficiency. Vortex to mix thoroughly and wait for 5 minutes.
[0117] 6.2 Briefly centrifuge and place on a magnetic rack for 2 minutes;
[0118] 6.3 Pipette the supernatant into a clean new centrifuge tube to collect the DNA solution. The purified DNA solution should be stored at -20℃.
[0119] (4) Multiplex PCR-NGS detection
[0120] In the first round of PCR, 194 sample nucleic acids were subjected to PCR reactions using methylation-specific primers (primer sequences and physical locations of methylation sites are shown in Table 1 ).
[0121] Table 1 Methylation sites
[0122] The reaction system for the first round of PCR is shown in Table 2. The PCR reaction conditions were: ① 95°C for 10 min; ② 35 cycles of the following: 95°C for 30 s; 48°C for 30 s; and 72°C for 30 s; and ③ 72°C for 5 min.
[0123] Table 2 Reaction system for the first round of PCR
[0124] For the second round of PCR, the 30 μL reaction system was as shown in Table 3.
[0125] AP5 primer sequence: AATGATACGGCGACCACCGAGATCTACACTCTTTCCCTACACGACGCTCTTCCGATCT (SEQ ID No. 21).
[0126] Index primer sequence: CAAGCAGAAGACGGCATACGAGATNNNNNNNNGTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT (SEQ ID No. 22), where NNNNNNNN represents an index, which is used to distinguish different samples.
[0127] The reaction conditions for the second round of PCR were as follows: ① 95°C for 10 min; ② 21 cycles of the following reactions: 95°C for 30 s; 57°C for 30 s; 72°C for 30 s; ③ 72°C for 5 min.
[0128] Table 3 Reaction system for the second round of PCR
[0129] The amplified product was purified by nucleic acid purification reagent to obtain a sequencing library, and then the sequencing reagent was used to TM Mid Output Reagent Cartridge (Illumina), sequencing was performed on a MiniSeq sequencer (Illumina).
[0130] (5) Calculation of methylation rate at each site
[0131] Statistical analysis was performed on the NGS results for 79 sites, with each methylation site sequenced at a depth of at least 500X. The number of reads containing a C base at a site was set to NumC, and the number of reads containing a T base at the same site was set to NumT. The ratio NumC / (NumC + NumT) was then used as the methylation rate for that site, i.e., the methylation data.
[0132] Example 2 Prediction Model Acquisition Process
[0133] (1) Dataset classification prediction
[0134] This specification example targets prostate cancer, benign prostatic hyperplasia, and other non-prostate cancer diseases, and includes a total of 194 samples from Example 1. Based on the combined PSA test results and targeted gene region methylation data, a process for obtaining and predicting a model for identifying prostate cancer samples is proposed, as shown in Figure 1.
[0135] Clinical samples were classified into high (H), medium (M), and low (L) risk categories based on fPSA / tPSA and tPSA thresholds, as shown in Table 4. The predictive classification algorithm consists of three modules, which classify clinical samples into six subcategories: high-risk prostate cancer (HC), medium-risk prostate cancer (MC), low-risk prostate cancer (LC), high-risk non-prostate cancer (HN), medium-risk non-prostate cancer (MN), and low-risk non-prostate cancer (LN). These subcategories serve as classification labels. A random forest algorithm was used to extract features from the methylation data for these six categories. The extracted features were then used to train a KNN model (left column of Figure 1) and predict the six categories for the prediction set. Support Vector Machines (SVM) were used to learn and predict these six features (right column of Figure 1). The two middle columns in Figure 1 show a Bayesian classifier that combines the KNN and SVM classification algorithms. A random forest algorithm is used to extract features from methylation sites that contribute significantly to the classification of clinical sample risk levels (high (H), medium (M), and low (L)). The KNN algorithm is then used to learn the training set after feature extraction and to make predictions on the prediction set, with the prediction results used as Bayesian likelihood probabilities. A random forest algorithm is then used to extract features from methylation sites in the training set that contribute significantly to the prostate cancer and non-prostate cancer classifications. The SVM algorithm is then used to learn the training set after feature extraction and to make predictions on the prediction set, with the prediction results used as Bayesian prior probabilities. Finally, the derived Bayesian formula is used to predict the risk level of clinical samples and the classification of prostate cancer and non-prostate cancer diseases.
[0136] Table 4 tPSA and fPSA thresholds for prostate cancer risk
[0137] (2) Model iterative optimization
[0138] The dataset and corresponding classification information were divided according to Tables 4 and 5, and the corresponding dataset was split into a training set and a prediction set. The model was constructed according to the process in Figure 1, the training set data features were learned, and predictions were made on the prediction set data. The model hyperparameters were solved according to the analysis process shown in Figure 2. During the iterative solution of the model hyperparameters, the iterations were terminated and the final prediction results were output when the four indicators of the prediction results given by the comprehensive evaluation model (the three evaluators' prediction of the six-category correlation prediction, the Bayesian likelihood probability risk level prediction, the three classifiers' prediction of prostate cancer and non-prostate cancer classification, and the consistency of the three classification evaluations) reached the set thresholds or the number of iterations reached 1000. The thresholds set here were: 1. The accuracy of the three evaluators' prediction of the six-category correlation prediction and the three classifiers' prediction of prostate cancer and non-prostate cancer classification reached 0.90; 2. The consistency of the three classification evaluations reached 0.90; and 3. The Bayesian likelihood probability risk level prediction reached 0.85.
[0139] Table 5 Clinical sample information and PSA classification and clinical classification information
[0140] During the model hyperparameter solution process, the number of model feature extractions and the model hyperparameter search range are appropriately adjusted. GridSearchCV is then used to search for the optimal estimator within the model hyperparameter search range. This estimator is used to learn the features of the training set data and predict the prediction set data. The optimal number of feature extractions and model hyperparameter ranges provided in the examples of this specification are shown in Table 6.
[0141] Table 6 Model hyperparameter search range
[0142] Example 3 Data Analysis and Results
[0143] (1) Sample feature extraction
[0144] A total of 250 clinically collected samples of prostate-related diseases were screened. After excluding samples without PSA data, 194 samples remained, of which 89 were clinically diagnosed as prostate cancer and 105 were diagnosed as non-prostate cancer, including 3 renal pelvis cancers, 6 bladder cancers, 3 renal malignancies and 3 prostate cysts, which served as interference samples for the accuracy of experimental analysis.
[0145] 1.1 Classification of clinical information
[0146] Among the collected clinical samples, most non-prostate cancer disease diagnoses include ureteral stones, benign prostatic hyperplasia, and other urinary tract inflammations. Prostate cancer diagnoses include prostate cancer, acinar prostate cancer, prostate cancer bone metastasis, and multiple prostate cancer metastases.
[0147] Based on the fPSA and tPSA values, clinical samples were divided into three risk levels: high (H), medium (M), and low (L). Combined with the clinical diagnosis of prostate cancer and non-prostate cancer results, six categories were finally formed. The methylation profile information of these clinical samples was reduced to two dimensions using UMAP (Uniform Manifold Approximation and Projection) technology, and a scatter plot was drawn. Color was used to distinguish prostate cancer (red) from non-prostate cancer (green), and shape was used to distinguish risk levels. The final six-category clustering is shown in Figures 3A and 3B.
[0148] As can be seen from Figures 3A and 3B, the methylation profiles of the samples exhibit a relatively obvious clustering phenomenon. The methylation sites used in the embodiments of this specification can be used to screen and differentiate between prostate cancer and non-prostate cancer diseases, and to a certain extent improve the ability to distinguish PSA risk levels.
[0149] 1.2 Samples are divided into training set and prediction set
[0150] The sample set of 194 cases (89 positive sample sets and 105 negative sample sets) was split into training set and prediction set in a ratio of 7:3, resulting in 135 randomly split training set samples and 59 prediction set samples. There were 4 interference samples in the prediction set samples, 1 prostate cyst and 3 bladder cancers, which served as interference samples for experimental analysis accuracy.
[0151] 1.3 Methylation feature extraction
[0152] 1.3.1 Methylation Feature Extraction of Prostate Cancer and Non-Prostate Cancer Samples
[0153] The random forest algorithm was used to calculate the feature contribution of methylation sites in the training set for distinguishing prostate cancer from non-prostate cancer. After screening the methylation sites, the site contribution was recalculated and a heat map was plotted, as shown in Figure 4. The Exp boxplot represents the overall methylation level of the sample; the longitudinal Exp represents the methylation level of the methylation site in the population; and the longitudinal Imp represents the random forest feature contribution level.
[0154] The figure shows the overall methylation level differences of the samples in the horizontal box plot and the methylation level differences of the methylation sites in the population in the vertical box plot. The selected methylation sites were used for classification using the Bayes classifier.
[0155] 1.3.2 Risk Level Feature Extraction
[0156] The random forest algorithm was used to calculate the feature contribution of methylation sites in the training set samples to distinguish the PSA-classified sample risk levels (H, M, and L). Vertical histograms of the feature contribution values of each methylation site were added, as shown in Figures 5A and 5B. The Exp boxplot represents the overall methylation level of the sample; the longitudinal Exp represents the methylation level of the methylation site in the population; and the longitudinal Imp represents the random forest feature contribution level.
[0157] As can be seen from the figure, the horizontal box plot shows the overall methylation level differences of the samples; the vertical methylation site methylation level differences in the population. The selected methylation sites were used for Bayesian risk prediction.
[0158] The 14 characteristic methylation sites extracted from prostate cancer and non-prostate cancer were combined with the 14 methylation sites extracted from risk levels for Bayesian prediction analysis. The combined characteristic 28 methylation sites are shown in Table 7.
[0159] Table 7 Random forest analysis of cancer vs. non-cancer and risk level feature extraction of methylation sites (28)
[0160] 1.3.3 Six-category feature extraction
[0161] The random forest algorithm was used to calculate the feature contribution values of the methylation sites in the training set to distinguish the six subtypes. By screening the methylation sites, the site contribution values were recalculated, and heat maps were drawn, as shown in Figures 6A and 6B.
[0162] As can be seen from the figure, the overall methylation level differences of the samples are shown in the horizontal box plot; the methylation level differences of the methylation sites in the population are shown in the vertical methylation sites. The selected methylation sites were used for training and prediction of the SVM six-classification evaluator and the KNN six-classification evaluator. The 28 methylation sites shown in Figures 6A and 6B are:
[0163] Table 8 Random Forest six-class feature extraction methylation sites (28)
[0164] (2) Classification performance
[0165] The sample set classification involved in the embodiments of this specification involves the following two types, including:
[0166] 1. Use PSA to classify samples into high risk (H), medium risk (M), and low risk (L). Use the fPSA threshold to classify the sample set based on the fPSA / tPSA value provided by the clinical sample information, and use the tPSA threshold to classify the sample set based on the tPSA value provided by the clinical sample information;
[0167] 2. The final clinical diagnosis of prostate cancer (C) and non-prostate cancer (N) is combined with the above risks to obtain: high-risk prostate cancer (HC), moderate-risk prostate cancer (MC), low-risk prostate cancer (LC), high-risk non-prostate cancer (HN), moderate-risk non-prostate cancer (MN) and low-risk non-prostate cancer (LN).
[0168] Based on the above sample set classification, the embodiments of this specification conduct a comprehensive analysis and evaluation from the following four aspects:
[0169] 2.1 Three evaluators predict six-category correlation evaluation
[0170] The KNN six-class estimator and the SVM six-class estimator can directly predict six categories: high-risk prostate cancer (HC), moderate-risk prostate cancer (MC), low-risk prostate cancer (LC), high-risk non-prostate cancer (HN), moderate-risk non-prostate cancer (MN), and low-risk non-prostate cancer (LN). The Bayesian estimator predicts the sample's risk level (high (H), moderate (M), low (L)) and prostate cancer (C) versus non-prostate cancer (N), and then uses the Bayesian formula to produce the final six-class prediction. The prediction performance of the three estimators after risk classification of the sample set using fPSA and tPSA is shown in Figure 7. The rows of the Confusion Matrix represent the six-class risk level prediction values of the three estimators, i.e., the prediction performance of the six-class risk level of the sample set using fPSA and tPSA. The columns of the matrix represent the label values of fPSA and tPSA, i.e., the six-class risk level of the sample based on the fPSA and tPSA provided by the clinical sample information. In the three evaluator modes, the numbers in the Confusion Matrix are mainly concentrated on the diagonal, indicating that the three evaluators have good prediction performance after risk classification of the sample set by fPSA and tPSA respectively.
[0171] 2.2 Bayesian likelihood probability risk level assessment
[0172] Risk assessment involves classifying samples into three subsets based on the PSA threshold: high (H), medium (M), and low (L). Methylation data recognition capabilities of these three subsets are then evaluated using the KNN, SVM, and Bayesian algorithms. Risk assessment is a merging of the prediction subsets from the six-category assessment. The specific operations are: 1. The risk assessment performance of the KNN and SVM classifiers is performed by extracting the six-category prediction results and merging the probabilities of the three risk levels being high (H), medium (M), and low (L); 2. The Bayesian risk assessment performance is performed by extracting the likelihood probabilities of the high (H), medium (M), and low (L) risk levels given by the KNN algorithm of the Bayesian estimator. The receiver operating characteristic (ROC) curves plotted using the extracted risk probabilities are shown in Figure 8, demonstrating good classification performance.
[0173] Specifically, the Bayesian likelihood probability risk assessment removes the interference of prostate cancer and non-prostate cancer classification performance and evaluates the risk level independently. The prediction results obtained by the Bayesian KNN algorithm are plotted in the confusion matrix shown in Figure 9, showing good prediction results.
[0174] 2.3 Evaluation of three classifiers for prostate cancer and non-prostate cancer classification
[0175] The Bayesian classification evaluator can directly predict whether a sample in the prediction set is prostate cancer or non-prostate cancer based on the Bayesian prior probability. The prediction results for prostate cancer and non-prostate cancer using the KNN and SVM six-class classification models are obtained by merging the six-class prediction results into two categories: prostate cancer and non-prostate cancer. The overall performance of the three classification evaluators for prostate cancer and non-prostate cancer prediction based on the fPSA and tPSA risk classification levels is shown in Figure 10, demonstrating strong predictive capabilities.
[0176] 2.4 Consistency evaluation of three classifiers
[0177] To better evaluate the consistency of the three evaluators, a comprehensive comparison of the prediction results for fPSA vs. tPSA was conducted on the prediction set samples. KNN, SVM, and the true labels were compared for prostate cancer vs. non-prostate cancer prediction. The consistency of the predictions for prostate cancer vs. non-prostate cancer is shown in the left column of Figure 11. For the six-class classification, a comprehensive comparison was made between the Bayesian evaluator, the SVM six-class evaluator, the KNN six-class evaluator, and the true clustering. The prediction consistency is shown in the Venn diagram in the right column of Figure 11, demonstrating good consistency.
[0178] Comprehensively evaluate the model's performance across the four dimensions mentioned above. When learning and predicting a sample set, only when the performance across the four dimensions reaches the expected threshold can model training and hyperparameter iteration be terminated, resulting in a trained, suitable classification prediction model suitable for clinical diagnosis.
[0179] (3) Comparison between model prediction and PSA interpretation
[0180] PSA detection and methylation profiles both play an important role in the identification of prostate cancer. PSA (prostate-specific antigen) is a protein produced by prostate cells. An increase in PSA levels can indicate the possible presence of prostate disease. Therefore, PSA detection is a commonly used method for screening prostate cancer. However, PSA detection cannot distinguish between prostate cancer and other prostate problems (such as benign hyperplasia), so further examination is needed to determine whether prostate cancer is present. Methylation profiles are descriptions of methylation patterns on genomic DNA. Methylation is a common form of epigenetic modification that can affect gene expression by changing the DNA methylation state. Prostate cancer generally has different methylation patterns compared to normal cells. Therefore, by analyzing the methylation profiles, characteristic markers or specific methylation patterns of prostate cancer can be found, thereby identifying prostate cancer. The predictive performance of the model constructed by the embodiments of this specification is compared with the PSA discrimination performance, as shown in Table 9.
[0181] Table 9 Comparison of PSA and model prediction of prostate cancer
[0182] Table 9 compares the risk levels predicted by the model, those predicted by PSA, and those interpreted clinically. The predictions from the three models are grouped into three risk categories: high, intermediate, and low. For example, the number 19 / 5 below the first H (high-risk group) in the fPSA group represents the number of subjects classified as high risk (H) based on fPSA, and the number of subjects classified as non-prostate cancer (N) by the clinical interpretation of the cell below. For example, the number 24 (0 / 3 / 21) in the SVM group represents the number of subjects classified as high risk, three intermediate risk, and 21 low risk, according to the SVM algorithm. Compared to the PSA predictions, the proportion of high-risk samples in the non-prostate cancer group has decreased significantly, while the proportion of low-risk samples has increased. This means that fewer patients are incorrectly diagnosed as high risk, thus reducing unnecessary further testing and treatment. At the same time, the proportion of low-risk samples in non-prostate cancer samples has increased, which means that more patients can be correctly diagnosed as low-risk, thus avoiding unnecessary anxiety and treatment. The 26 (4 / 22 / 0) after SVM indicates that, according to the algorithm prediction of the embodiment of this specification, among the 26 prostate samples, there are 4 high-risk, 22 medium-risk, and 0 low-risk. The proportion of low-risk samples in prostate cancer samples has been significantly reduced, reducing the probability of low-risk patients being mistakenly diagnosed with prostate cancer. This improvement in accuracy helps avoid unnecessary further examinations and treatments, reducing the burden and discomfort on patients. It can be clearly seen from the comparison table that the SVM, KNN, and Bayesian models constructed in the embodiment of this specification all have better ability to identify prostate cancer risk levels than PSA testing, and can more accurately classify samples. These models all have good sensitivity. As an example, in distinguishing prostate cancer samples, the sensitivity of the Bayesian algorithm can reach 28 / 31 = 90%. This improvement in effect has important clinical significance for the diagnosis and treatment of prostate cancer.
[0183] The above embodiment establishes a new method for early prostate cancer screening by combining methylation of the prostate-specific antigen (PSA) indicator. Using the widely used PSA threshold in clinical practice, clinical samples are divided into three categories: high-risk (H), medium-risk (M), and low-risk (L), and ultimately diagnosed as prostate cancer (C) and non-prostate cancer (N). Specifically, there are two types of early prostate cancer screening methods:
[0184] 1. Using the Support Vector Machine (SVM) algorithm to learn the 28 methylation site profiles extracted from the training set of prostate cancer (C), non-prostate cancer (N), and risk level (HML), the prediction results for prostate cancer (C) and non-prostate cancer (N) in the prediction set are obtained as Bayesian prior probabilities. The KNN algorithm is used to calculate the likelihood of the classification of the prediction samples based on the PSA threshold. Then, the Bayesian formula is used to infer the classification probabilities of the six categories of the prediction set samples and the probabilities of prostate cancer and non-prostate cancer.
[0185] 2. Based on the aforementioned prostate cancer and non-prostate cancer risk classification schemes, samples were divided into six categories: high-risk prostate cancer (HC), moderate-risk prostate cancer (MC), low-risk prostate cancer (LC), high-risk non-prostate cancer (HN), moderate-risk non-prostate cancer (MN), and low-risk non-prostate cancer (LN). KNN and SVM algorithms were used to learn the methylation profiles of the 28 training sets extracted by random forests. The classification probabilities of the six categories and the probabilities of prostate cancer and non-prostate cancer were predicted for the prediction set samples.
[0186] It should be noted that since most of the negative clinical samples are from patients with prostate diseases, such as benign prostatic hyperplasia, etc., the feasibility of PSA detection technology in distinguishing prostate cancer from non-prostate cancer is not high. Clinically, PSA testing is generally performed only when there is a prostate lesion to confirm whether there is a disease. Therefore, in actual operation, using PSA detection technology to distinguish prostate cancer from other prostate diseases (benign prostatic hyperplasia) is a technical problem that must be faced and solved. After the patient has completed the PSA test, if the PSA value is found to be abnormal, combined with the methylation detection and analysis strategy provided in the embodiment, the accuracy of PSA in prostate cancer screening can be significantly improved.
[0187] While the basic concepts have been described above, it will be apparent to those skilled in the art that the detailed disclosure is merely illustrative and does not limit this specification. Although not explicitly stated herein, various modifications, improvements, and revisions to this specification may be made by those skilled in the art. Such modifications, improvements, and revisions are suggested in this specification and remain within the spirit and scope of the exemplary embodiments of this specification.
[0188] This specification also uses specific terms to describe the embodiments of this specification. For example, "one embodiment," "an embodiment," and / or "some embodiments" refer to a feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "one embodiment," "an embodiment," or "an alternative embodiment" two or more times in different locations in this specification do not necessarily refer to the same embodiment. Furthermore, certain features, structures, or characteristics of one or more embodiments of this specification may be appropriately combined.
[0189] In addition, unless expressly stated in the claims, the order of the processing elements and sequences, the use of alphanumeric characters, or the use of other names described in this specification are not intended to limit the order of the processes and methods of this specification. Although the above disclosure discusses some of the invention embodiments currently considered useful through various examples, it should be understood that such details are for illustrative purposes only, and the appended claims are not limited to the disclosed embodiments. On the contrary, the claims are intended to cover all modifications and equivalent combinations that are consistent with the spirit and scope of the embodiments of this specification. For example, although the system components described above can be implemented by hardware devices, they can also be implemented only by software solutions, such as installing the described system on an existing server or mobile device.
[0190] Similarly, it should be noted that, in order to simplify the presentation of this specification and thus facilitate understanding of one or more embodiments of the invention, the foregoing descriptions of the embodiments of this specification sometimes combine multiple features into a single embodiment, figure, or description thereof. However, this disclosure method does not imply that the subject matter of this specification requires more features than those recited in the claims. In fact, an embodiment may have fewer features than all of the features of a single disclosed embodiment.
[0191] In some embodiments, numbers are used to describe the quantity of components and attributes. It should be understood that such numbers used in the description of the embodiments are modified by the modifiers "about", "approximately" or "substantially" in some examples. Unless otherwise stated, "about", "approximately" or "substantially" indicate that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the description and claims are approximate values, which may change according to the required characteristics of individual embodiments. In some embodiments, the numerical parameters should take into account the specified significant digits and adopt the general method of retaining digits. Although the numerical domains and parameters used to confirm the breadth of their range in some embodiments of this specification are approximate values, in specific embodiments, the settings of such numerical values are as accurate as possible within the feasible range.
[0192] Each patent, patent application, patent application publication, and other materials, such as articles, books, specifications, publications, and documents, cited in this specification is hereby incorporated by reference in its entirety. This includes application history documents that are inconsistent with or conflict with the content of this specification, as well as documents (currently or subsequently attached to this specification) that limit the broadest scope of the claims of this specification. It should be noted that if the descriptions, definitions, and / or terminology used in the accompanying materials are inconsistent or conflicting with the content of this specification, the descriptions, definitions, and / or terminology used in this specification will control.
[0193] Finally, it should be understood that the embodiments described in this specification are intended only to illustrate the principles of the embodiments of this specification. Other variations may also fall within the scope of this specification. Therefore, by way of example and not limitation, alternative configurations of the embodiments of this specification may be considered consistent with the teachings of this specification. Accordingly, the embodiments of this specification are not limited to the embodiments explicitly described and illustrated in this specification.
Claims
1. A method for generating a cancer risk level prediction model, wherein: The prediction model includes a first prediction model, a second prediction model, a third prediction model, a fourth prediction model and a fifth prediction model, and the method includes: Step S1, using a training set and labels to respectively train an initial first prediction model, an initial second prediction model, an initial third prediction model, and an initial fourth prediction model to obtain trained first prediction models, second prediction models, third prediction models, and fourth prediction models, wherein the training set includes methylation data of cancer-related methylation sites of each sample in the training samples, and the label is the cancer risk level of the sample; Step S2, respectively using the trained first prediction model, second prediction model, third prediction model, and fourth prediction model to predict samples in the prediction set, to obtain prediction results of the four prediction models, wherein the prediction results of the third prediction model and the fourth prediction model are input into the fifth prediction model to obtain the prediction result of the fifth prediction model; and Step S3, determining whether each prediction result meets the preset conditions, if so, generating each prediction model; if not, returning to step S1 for further iteration until each prediction result meets the preset conditions.
2. The method according to claim 1, wherein The first and third prediction models are KNN models, the second and fourth prediction models are SVM models, and the fifth prediction model is a Bayesian classifier.
3. The method according to claim 1 or 2, wherein: The first and second prediction models use the same labels and the same training set data; the third and fourth prediction models use different labels but the same, different or overlapping training set data.
4. The method according to claim 1 or 2, wherein: The risk levels of the samples having the cancer in the labels of the first and second prediction models are high probability of cancer, medium probability of cancer, low probability of cancer, high probability of non-cancer, medium probability of non-cancer and low probability of non-cancer.
5. The method according to claim 1 or 2, wherein: The risk levels of the sample having the cancer in the label of the third prediction model are high probability, medium probability and low probability; The fourth prediction model labels the sample as having the cancer risk level of cancer or non-cancer.
6. The method according to claim 5, wherein The risk levels of the sample having the cancer in the labels of the third prediction model as high probability, medium probability, and low probability are determined based on the values of the cancer markers in the sample.
7. The method according to claim 6, wherein The cancer is prostate cancer; and the cancer marker is PSA.
8. The method according to claim 6 or 7, wherein: When the value of the cancer marker is tPSA, tPSA is less than the first concentration value, and the risk level of the sample having the cancer is low; tPSA is between the first concentration value and the second concentration value, and the risk level of the sample having the cancer is medium; If tPSA is greater than the second concentration value, the risk level of the sample having the cancer is high.
9. The method according to claim 6 or 7, wherein: When the value of the cancer marker is fPSA / tPSA, fPSA / tPSA is greater than a first value, indicating that the sample has a low risk of having the cancer; When fPSA / tPSA is between the first value and the second value, the risk level of the sample suffering from the cancer is medium; If fPSA / tPSA is less than the second value, the risk level of the sample having the cancer is high.
10. The method according to claim 4, wherein The risk levels of the sample having the cancer in the labels of the first and second prediction models are high chance of cancer, medium chance of cancer, low chance of cancer, high chance of non-cancer, medium chance of non-cancer, and low chance of non-cancer, which are determined based on the values of the cancer markers in the sample and the clinical classification.
11. The method according to claim 1 or 2, wherein: The preset conditions include: The accuracy of the prediction results of the first prediction model, the second prediction model, and the fifth prediction model respectively reaches a first threshold; Converting the prediction results of the first prediction model and the second prediction model into corresponding prediction results of the third prediction model, wherein the accuracy of the converted prediction results and the prediction results of the third prediction model respectively reaches a second threshold; Converting the prediction results of the first prediction model and the second prediction model into prediction results corresponding to the prediction results of the fourth prediction model, wherein the accuracy of the converted prediction results and the prediction results of the fourth prediction model respectively reaches a third threshold; and The consistency of the prediction results of the first prediction model, the second prediction model and the fifth prediction model reaches a fourth threshold.
12. The method according to claim 1 or 2, wherein: The preset condition includes: the number of iterations reaches a preset iteration value.
13. The method according to claim 1 or 2, wherein: The methylation sites in the training set used to train different prediction models were screened using a classification algorithm.
14. The method according to claim 13, wherein The classification algorithm is a random forest algorithm.
15. The method according to claim 8, wherein The methylation sites used to train the first prediction model and the second prediction model include at least: IGFBP3_83, IGFBP3_16, IGFBP3_121, IGFBP3_48, IGFBP3_36, FEZF2_21, IGFBP3_44, IGFBP3_114, SERPINB1_4, FEZF2_74, FEZF2_5, IGFBP3_67, SERPINB1_67, SERPINB1_ 97, ZNF154_6, SOX1-OT_5, FEZF2_23, APC_78, FHAD1_57, SERPINB1_51, APC_98, IGFB P3_95, MIR663A_3, inTerGrnTiC_1_84, SERPINB1_36, POU4F2_77, APC_36 and APC_110.
16. The method according to claim 9, wherein The methylation sites used to train the first prediction model and the second prediction model include at least: IGFBP3_36, IGFBP3_121, SERPINB1_4, IGFBP3_89, FEZF2_5, SOX1-OT_3, MIR663A_55, FH AD1_27, FEZF2_74, SOX1-OT_58, MIR663A_61, MIR663A_90, APC_105, POU4F2_77, POU4F2 _98, FHAD1_2, MIR663A_50, inTerGrnTiC_1_69, APC_36, FHAD1_57, inTerGrnTiC_1_53, SOX1-OT_5, SOX1-OT_74, SERPINB1_91, IGFBP3_67, APC_117, IGFBP3_101 and ZNF154_116.
17. The method according to claim 8, wherein The methylation sites used to train the third prediction model include at least: SERPINB1_4, IGFBP3_36, IGFBP3_121, IGFBP3_114, SERPINB1_67, FEZF2_21, IGFBP3_16, FHAD1_2, IGFBP3_101, FEZF2_31, SOX1-OT_105, IGFBP3_83, FEZF2_74, and IGFBP3_44.
18. The method according to claim 9, wherein The methylation sites used to train the third prediction model include at least: IGFBP3_114, MIR663A_39, IGFBP3_16, ZNF154_95, SOX1-OT_89, FHAD1_27, APC_20, IGFBP3_44, FEZF2_23, IGFBP3_48, SOX1-OT_74, SOX1-OT_3, IGFBP3_36 and SOX1-OT_5.
19. The method according to claim 1 or 2, wherein: The methylation sites used to train the fourth prediction model include at least: FEZF2_21, IGFBP3_121, IGFBP3_114, MIR663A_3, IGFBP3_36, FHAD1_2, ZNF154_116, SOX1-OT_58, IGFBP3_16, FEZF2_66, SERPINB1_4, SERPINB1_34, APC_78 and IGFBP3_101.
20. A system for determining a prediction model capable of predicting a subject's risk level of developing cancer, characterized in that: include: The training module is used to perform the following operations: Step S1, using a training set and labels to respectively train an initial first prediction model, an initial second prediction model, an initial third prediction model, and an initial fourth prediction model to obtain trained first prediction models, second prediction models, third prediction models, and fourth prediction models, wherein the training set includes methylation data of cancer-related methylation sites of each sample in the training samples, and the label is the cancer risk level of the sample; Step S2, using the trained first prediction model, second prediction model, third prediction model, and fourth prediction model to predict samples in the prediction set, respectively, to obtain prediction results of the four prediction models, wherein the prediction results of the third prediction model and the fourth prediction model are input into the fifth prediction model to obtain the prediction result of the fifth prediction model; as well as Step S3, determining whether each of the prediction results meets the preset conditions, if so, generating each prediction model; if not, returning to step S1 for further iteration until each prediction result meets the preset conditions.
21. A device for determining a prediction model capable of predicting a subject's risk level of developing cancer, characterized in that: The apparatus comprises at least one processor and at least one memory; The at least one memory is for storing computer instructions; The at least one processor is configured to execute at least part of the computer instructions to implement the method according to any one of claims 1 to 19.
22. A computer-readable storage medium, characterized in that The storage medium stores computer instructions. When a computer reads the computer instructions in the storage medium, the computer executes the method according to any one of claims 1 to 19.
23. A method for predicting a subject's risk level of developing cancer, characterized in that: The method comprises: Obtaining methylation data for methylation sites associated with the cancer; and The methylation data of the subject is input into a prediction model to predict the risk level of the subject suffering from cancer, wherein the prediction model is a first prediction model, a second prediction model, a third prediction model, a fourth prediction model, or a fifth prediction model, wherein the prediction model is generated by: Step S1, using a training set and labels to respectively train an initial first prediction model, an initial second prediction model, an initial third prediction model, and an initial fourth prediction model to obtain trained first prediction models, second prediction models, third prediction models, and fourth prediction models, wherein the training set includes methylation data of cancer-related methylation sites of each sample in the training samples, and the label is the cancer risk level of the sample; Step S2, respectively using the trained first prediction model, second prediction model, third prediction model, and fourth prediction model to predict samples in the prediction set, to obtain prediction results of the four prediction models, wherein the prediction results of the third prediction model and the fourth prediction model are input into the fifth prediction model to obtain the prediction result of the fifth prediction model; and Step S3, determining whether each of the prediction results meets the preset conditions, if so, generating each prediction model; if not, returning to step S1 for further iteration until each prediction result meets the preset conditions.
24. A system for predicting a subject's risk level of having cancer, characterized in that: The system comprises: an acquisition module, for acquiring methylation data of methylation sites associated with the cancer; A prediction module is configured to input the methylation data of the subject into a prediction model to predict the risk level of the subject suffering from cancer, wherein the prediction model is generated by: Step S1, using a training set and labels to respectively train an initial first prediction model, an initial second prediction model, an initial third prediction model, and an initial fourth prediction model to obtain trained first prediction models, second prediction models, third prediction models, and fourth prediction models, wherein the training set includes methylation data of cancer-related methylation sites of each sample in the training samples, and the label is the cancer risk level of the sample; Step S2, respectively using the trained first prediction model, the second prediction model, the third prediction model and the fourth prediction model to predict the samples in the prediction set, and obtaining the prediction results of the four prediction models, wherein the third prediction model and the fourth prediction model are used to predict the samples in the prediction set. The prediction result of the prediction model is input into the fifth prediction model to obtain the prediction result of the fifth prediction model; and Step S3, judging whether each of the prediction results meets the preset conditions, if so, the training is completed; if not, returning to step S1 for further iteration until each prediction result meets the preset conditions.
25. A device for predicting a subject's risk level of having cancer, characterized in that: The apparatus comprises at least one processor and at least one memory; The at least one memory is for storing computer instructions; The at least one processor is configured to execute at least part of the computer instructions to implement the method of claim 23.
26. A computer-readable storage medium, characterized in that The storage medium stores computer instructions. When a computer reads the computer instructions in the storage medium, the computer executes the method according to claim 23.
Citation Information
Patent Citations
Early non-small cell lung cancer recurrence model construction method based on DNA methylation
CN111564177A
Colorectal cancer risk prediction method and system, computer equipment and readable storage medium
CN111739642A
Lung cancer diagnosis system based on multiple machine learning algorithms
CN112259221A
Pulmonary nodule classification method and product based on lung CT and polygene methylation
CN115984251A
Lung cancer risk prediction method based on machine learning and related equipment
CN117542515A
Cited By
Fall risk detection method based on multi-modal health data
CN121682501A