Combinations of methylation biomarkers for multiple cancer types, screening methods and their applications
By screening cancer tissue and body fluid samples through methylation sequencing, a combination of low-noise and high-sensitivity methylation biomarkers was constructed, solving the problems of high false positive rates and biomarker redundancy in multiple cancer screening. This enabled early screening and accurate localization of multiple cancers, reduced the medical burden, and improved patient survival rates.
Patent Information
- Application Number
- CN202410881225.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-02
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2044-07-02
AI Technical Summary
Existing technologies suffer from high false positive rates, redundant biomarkers, and low screening efficiency in multi-cancer screening, especially for healthy and high-risk individuals, making it difficult to achieve efficient and accurate screening for multiple cancers.
By sequencing the methylation of cancerous and non-cancer tissues, differentially methylated regions are screened out. Combined with targeted methylation sequencing of body fluid samples, combinations of low-noise and high-sensitivity methylation biomarkers are screened out. The biomarker combinations are then optimized through iterative algorithms to construct a multi-cancer early screening and source tracing model.
It achieves highly sensitive and low false-positive screening for various cancers, enabling early detection of cancer and accurate location of cancer sites, reducing medical burden, and improving patient survival rate and screening efficiency.
Smart Images

Figure CN118755832B_ABST
Abstract
Description
Technical Field
[0001] This disclosure belongs to the biomedical field and mainly relates to a screening method for combinations of methylation markers for multiple cancer types, the screened combinations of methylation markers and their applications. Background Technology
[0002] Cancer infection leads to systemic organ failure, ultimately causing death. Cancer poses a significant threat to global public health. Most patients are diagnosed at middle or late stages, resulting in poor treatment outcomes, shortened survival, and a heavy economic burden. Studies show that early screening and intervention can effectively improve patients' quality of life and increase five-year survival rates, and also has significant socioeconomic benefits. Therefore, early cancer screening is of great importance.
[0003] Currently, single-cancer screening is a common clinical method, suitable for auxiliary diagnosis in high-risk groups, and has significant effects on the early detection and treatment of specific cancers. Some cancer risk factors (such as alcoholism and smoking) are associated with multiple cancers. When using single-cancer screening to screen healthy and high-risk individuals, the cumulative false-positive rate due to multiple tests may increase the medical burden. Multi-cancer screening, on the other hand, can simultaneously screen and assist in the diagnosis of multiple cancers in asymptomatic healthy individuals, and conduct cancer risk assessments. It has great potential, especially for those with a family history, high environmental exposure risk, or known genetic susceptibility factors, improving screening efficiency and expanding the screening scope. Therefore, developing a simple, accurate, and non-invasive method for the early detection of multiple cancers is of great significance.
[0004] DNA methylation sequencing, as a high-resolution, high-throughput technology, is applied to cancer screening, diagnosis, and monitoring. Currently, most ctDNA methylation biomarker studies rely on tissue methylation microarray results or tissue whole-genome DNA methylation sequencing (WGBS) results. The former has the limitation of missing some high-performance biomarkers, while using only the latter results in a large set of biomarkers to be selected and retained, and biomarker redundancy. Summary of the Invention
[0005] To address at least one of the above problems, this disclosure provides a method for screening combinations of multiple cancer methylation biomarkers, as well as the multiple cancer methylation biomarkers obtained by this method and their applications.
[0006] According to one aspect of this disclosure, a method for screening methylation markers is provided, the method comprising:
[0007] S1) Obtain the methylation sequencing results of cancer tissue samples and non-cancer samples corresponding to the cancer type, and divide the genome into multiple methylation regions based on the methylation sequencing results;
[0008] S2) Perform differential analysis on the methylation regions of cancer tissue samples and non-cancer samples of the cancer type, screen out the methylation regions whose average methylation rate difference of the cancer type is greater than a threshold fraction relative to non-cancer samples, and merge and remove duplicates to obtain the biomarker set A;
[0009] S3) By performing differential analysis on the methylation regions of the cancer tissue samples of the stated cancer type and tissue samples of one or more other cancer types, methylation regions with an average methylation rate difference greater than a threshold fraction relative to tissue samples of one or more other cancer types are screened out, and after merging and deduplication, a set of markers B is obtained.
[0010] S4) Targeted methylation sequencing is performed on biomarker set A in body fluid samples and non-cancer body fluid samples of the cancer type to obtain the coverage E1 of each biomarker in the cancer type, and biomarker set C is obtained based on the coverage E1.
[0011] S5) Targeted methylation sequencing is performed on biomarker set B in body fluid samples of the stated cancer type and one or more other cancer types to obtain the coverage E2 of each biomarker in the stated cancer type, and biomarker set D is obtained based on the coverage E2; and
[0012] S6) Optionally, merge the set of markers C and the set of markers D to remove duplicates and obtain the set of markers.
[0013] In some embodiments, the division of the genome into multiple methylation regions in step S1) can be performed using methods conventionally used by those skilled in the art.
[0014] In some embodiments, in step S3), a differential analysis is performed on the methylation regions of the cancer tissue sample of the stated cancer type compared to tissue samples of another cancer type. In optional embodiments, in step S3), a differential analysis is performed on the methylation regions of the cancer tissue sample of the stated cancer type compared to tissue samples of multiple other cancer types, either individually or simultaneously.
[0015] In some embodiments, in steps S2) and S3), screening out methylation regions with an average methylation rate difference greater than a threshold fraction for the cancer type means screening out methylation regions with significant differences, preferably screening out methylation regions with significantly high methylation rates.
[0016] The difference analysis can be performed using any significance statistical method known to those skilled in the art, such as rank-sum test, T-test, analysis of variance, etc.
[0017] In a specific implementation, in steps S2) and S3), the screening of methylated regions where the average methylation rate difference of the cancer type is greater than a threshold fraction includes: screening methylated regions where the average methylation rate difference of the cancer tissue sample is greater than 0.1 (e.g., greater than 0.1, greater than 0.12, greater than 0.13, greater than 0.14, greater than 0.15, greater than 0.16, greater than 0.17, greater than 0.18, greater than 0.19, or greater than 0.2), and / or where the P value is less than 0.5 (e.g., less than 0.4, less than 0.3, less than 0.2, less than 0.1, less than 0.09, less than 0.08, less than 0.07, less than 0.06, less than 0.05, less than 0.04, less than 0.03, less than 0.02, or less than 0.01).
[0018] In some embodiments, the average methylation rate refers to the proportion of the number of methylated CpG sites on all reads within a marker to the total number of CpG sites detected in that marker.
[0019] In some implementations, step S2) and / or step S3) may also include filtering the set of markers A and / or the set of markers B.
[0020] In some embodiments, the filtering step includes: selecting a set A' of markers from marker set A that has an average methylation rate of less than 0.5 (e.g., 0.5, 0.4, 0.3, 0.2, or 0.1) in non-cancerous body fluid samples. In a preferred embodiment, the filtering step includes: selecting a set A' of markers from marker set A that has an average methylation rate of less than 0.5 (e.g., 0.5, 0.4, 0.3, 0.2, or 0.1) in regions at the 85th to 99.9th percentile (preferably the 95th percentile) in non-cancerous body fluid samples.
[0021] In some embodiments, the filtering step includes: selecting from the biomarker set B a set B' of biomarkers with an average methylation rate below 0.5 (e.g., 0.5, 0.4, 0.3, 0.2, or 0.1) in tissue samples of one or more other cancer types. In a preferred embodiment, the filtering step includes: selecting from the biomarker set B a set B' of biomarkers with an average methylation rate below 0.5 (e.g., 0.5, 0.4, 0.3, 0.2, or 0.1) in regions at the 85th to 99.9th percentile (preferably the 95th percentile) in tissue samples of one or more other cancer types.
[0022] In some embodiments, step S4) involves targeted methylation sequencing of the biomarker set A' in the body fluid sample and non-cancer body fluid sample of the cancer type to obtain the coverage E1 of each biomarker in the cancer type, and obtaining the biomarker set C based on the coverage E1.
[0023] In some embodiments, step S5) involves targeted methylation sequencing of a biomarker set B' in a body fluid sample of the cancer type and a body fluid sample of one or more other cancer types to obtain the coverage E2 of each biomarker in the cancer type, and obtaining a biomarker set D based on the coverage E2.
[0024] In some embodiments, step S4) of obtaining the coverage E1 of each marker in the cancer type includes: selecting reads with a methylation rate of 0.5 or higher (e.g., 0.6, 0.7, or 0.8 or higher) and covering more than 3 CpG sites as extremely hypermethylated reads, the proportion of an extremely hypermethylated read of a marker to the total reads of that marker as the extreme hypermethylated read ratio of that marker, and binary processing the obtained extreme hypermethylated read ratios of each marker in the cancer type body fluid sample based on threshold quantiles to determine the coverage E1 of each marker.
[0025] In some embodiments, in step S4), the threshold quantile can be determined by statistically analyzing the extreme hypermethylation read ratios in non-cancerous blood samples. In some embodiments, determining the threshold quantile in step S4) includes: statistically analyzing the extreme hypermethylation read ratios of each biomarker in non-cancerous blood samples, and using the 90%–99% quantiles of these extreme hypermethylation read ratios (e.g., 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% quantiles) as the threshold quantiles for each biomarker.
[0026] In some implementations, step S4) includes the binarying process as follows: if the extreme hypermethylation read ratio of a marker in a cancer blood sample is higher than the threshold quantile, then the marker is set to 1, otherwise it is set to 0, and vice versa.
[0027] In some embodiments, step S4), obtaining the marker set C based on coverage E1 includes selecting the top 200 markers by coverage and / or markers with coverage greater than 1%. In some embodiments, step S4), selecting markers with coverage ranking in the top 256, top 128, top 64, top 32, top 16, top 8, top 4, or top 2, and / or markers with coverage greater than 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10%.
[0028] In some embodiments, step S5) of obtaining the coverage E2 of each marker in the cancer type includes: selecting reads with a methylation rate of 0.5 or higher (e.g., 0.6, 0.7, or 0.8 or higher) and covering more than 3 CpG sites as extremely hypermethylated reads, the proportion of an extremely hypermethylated read of a marker to the total reads of that marker as the extreme hypermethylated read ratio of that marker, and binary processing the obtained extreme hypermethylated read ratios of each marker in the cancer type body fluid sample based on threshold quantiles to determine the coverage E2 of each marker.
[0029] In some embodiments, the methylation rate includes the proportion of the number of methylated CpG sites to the total number of CpG sites in the read.
[0030] In some embodiments, in step S5), the threshold quantile can be determined by statistically analyzing the extreme hypermethylated read ratios in body fluid samples from one or more other cancer types. In some embodiments, determining the threshold quantile in step S5) includes: statistically analyzing the extreme hypermethylated read ratios of each marker in body fluid samples from one or more other cancer types, and using the 90% to 99% quantiles of the extreme hypermethylated read ratios (e.g., 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% quantiles) as the threshold quantiles for each marker.
[0031] In some implementations, step S5) may include setting the binary process to 1 if the extreme hypermethylation read ratio of a marker in a cancer blood sample is higher than the threshold quantile, otherwise setting it to 0, and vice versa.
[0032] In some implementations, step S5), obtaining the marker set D based on coverage E2, includes selecting the top 200 markers by coverage and / or markers with coverage greater than 1%. In some implementations, step S5), selecting markers with coverage ranking in the top 256, top 128, top 64, top 32, top 16, top 8, top 4, or top 2, and / or markers with coverage greater than 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10%.
[0033] In some embodiments, the method may further include: using a body fluid sample of the cancer type and the non-cancer body fluid sample, employing an iterative algorithm to screen for sensitive methylation biomarkers from biomarker set A and / or biomarker set A', and adding them to biomarker set C. In some embodiments, the step of using an iterative algorithm to screen for sensitive methylation biomarkers from biomarker set A and / or biomarker set A' includes:
[0034] M1) The complementarity gain score Qi for each newly added marker i is calculated using the following formula:
[0035]
[0036] Where the weight coefficient W1 is 1, the weight coefficient W2 is -1, the weight coefficient W3 is -2, and N C For cancer-specific body fluid samples in the cancer-specific biomarker cohort of biomarker set A and / or biomarker set A', where the proportion of positive biomarkers is below the threshold P1, N h1 For non-cancerous body fluid samples in the biomarker set A and / or biomarker set A', where the proportion of positive biomarkers is higher than threshold P2 but lower than threshold P3, N h2 For non-cancerous body fluid samples in the biomarker set A and / or biomarker set A', where the proportion of positive biomarkers is higher than the threshold P3, R in The proportion of extremely high methylation reads representing the new biomarker i in n samples, C i The threshold quantile representing the newly added marker i,
[0037] The positive biomarker is a biomarker with an extremely high methylation read ratio higher than the threshold quantile;
[0038] Threshold P1 is the 50%–70% quantile (e.g., 50%, 60%, or 70%) of the proportion of positive markers in cancerous body fluid samples; threshold P2 is the 85%–93% quantile (e.g., 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, or 93%) of the proportion of positive markers in non-cancer body fluid samples; threshold P3 is the 93%–96% quantile (e.g., 93%, 94%, 95%, or 96%) of the proportion of positive markers in non-cancer body fluid samples.
[0039] M2) Based on the complementary gain score Qi, the marker with the highest score in each iteration is selected as the new sensitivity marker; and
[0040] Optionally, in the M3 process, two conditions are set as iteration termination conditions: ① When Qi is less than or equal to 0, the iteration process of adding new biomarkers is stopped; ② Using the 90% to 99% quantile (e.g., 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% quantile) of the proportion of positive biomarkers in non-cancer samples as a threshold, the iteration process is stopped when the number of cancer body fluid samples above the threshold after adding the new biomarker is less than the number of cancer body fluid samples above the threshold before adding the new biomarker.
[0041] In the method disclosed herein, there is no particular limitation on the type of cancer, including any type of cancer.
[0042] In a specific implementation, the cancer type is selected from any one of liver cancer, lung cancer, colorectal cancer, stomach cancer, esophageal cancer, pancreatic cancer, breast cancer, and ovarian cancer.
[0043] In some implementations, the non-cancerous sample includes adjacent tissue or body fluid samples.
[0044] In some embodiments, the body fluid samples include: saliva or saliva supernatant, whole blood, serum, plasma, milk, urine, lumbar or ventricular CSF, lymph, prostatic fluid, semen, sputum, diluted fecal supernatant, tears, tumor cells, bronchoalveolar lavage fluid, sputum, pus, nasopharyngeal swabs, oral swabs, cerebrospinal fluid, pleural and peritoneal fluid, amniotic fluid, peritoneal fluid, aqueous humor, vitreous humor, diluted vaginal discharge supernatant, and their processed forms.
[0045] In some embodiments, the tissue sample includes tissue, paraffin sections, and their processed forms.
[0046] According to another aspect of this disclosure, an apparatus for screening methylation markers is provided, which is used to implement the screening method for methylation markers described above.
[0047] In some embodiments, the apparatus for screening methylation markers includes:
[0048] A module for segmenting methylation regions; used to divide the genome into multiple methylation regions based on methylation sequencing results from cancer tissue samples and non-cancer samples corresponding to cancer types;
[0049] The first methylation region differential analysis module is used to perform differential analysis on the methylation regions of cancer tissue samples and non-cancer samples of the cancer type, screen out methylation regions whose average methylation rate difference of the cancer type is greater than a threshold fraction relative to non-cancer samples, and obtain the biomarker set A after merging and deduplication.
[0050] The second methylation region differential analysis module is used to perform differential analysis on the methylation regions of the cancer tissue sample of the stated cancer type and tissue samples of one or more other cancer types, screen out the methylation regions whose average methylation rate difference is greater than a threshold fraction relative to tissue samples of one or more other cancer types, and merge and remove duplicates to obtain the biomarker set B.
[0051] The first coverage analysis module is used to obtain the coverage E1 of each biomarker in the cancer type based on the targeted methylation sequencing results of biomarker set A in body fluid samples and non-cancer body fluid samples of the cancer type, and obtain biomarker set C based on coverage E1.
[0052] The second coverage analysis module is used to obtain the coverage E2 of each biomarker in the cancer type based on the targeted methylation sequencing results of biomarker set B in body fluid samples of the stated cancer type and one or more other cancer types, and to obtain biomarker set D based on coverage E2; and
[0053] Optionally, the marker set acquisition module is used to merge and deduplicate marker set C and marker set D to obtain the marker set.
[0054] In some embodiments, the methylation biomarker screening device further includes an iterative algorithm module that uses body fluid samples of the cancer type and the non-cancer body fluid samples to screen sensitive methylation biomarkers from biomarker set A and / or biomarker set A' and add them to biomarker set C.
[0055] According to another aspect of this disclosure, a combination of methylation markers is provided, obtained by the screening method for the methylation markers described above.
[0056] In some embodiments, the methylation biomarker combination includes cancer detection methylation biomarkers and / or organ-derived methylation biomarkers.
[0057] In some implementations, the cancer detection methylation biomarkers are based on the sequence of the human reference genome Hg19, and include any one or more of the methylation biomarkers shown in Table 1.
[0058] In some implementations, based on the sequence of the human reference genome Hg19, the organ-derived methylation markers include any one or more combinations of methylation markers shown in Table 2.
[0059] In some embodiments, the cancer types include one or more of liver cancer, lung cancer, colorectal cancer, stomach cancer, esophageal cancer, pancreatic cancer, breast cancer, and ovarian cancer.
[0060] According to another aspect of this disclosure, a primer pair is provided, wherein the primer pair uses the nucleotide sequence containing the methylation marker combination as the target sequence for specific amplification of the target sequence.
[0061] According to another aspect of this disclosure, a probe is provided for specifically capturing the methylation markers.
[0062] According to another aspect of this disclosure, a kit is provided for detecting the above-described combinations of methylation markers of this disclosure.
[0063] In some embodiments, the kit includes the primer pair and / or the probe.
[0064] In some embodiments, the kit further includes a conversion agent capable of converting methylated or unmethylated cytosine into uracil.
[0065] According to another aspect of this disclosure, a multi-cancer detection model is provided, which is constructed using the above-described combination of methylation markers of this disclosure.
[0066] According to another aspect of this disclosure, a method for detecting multiple cancer types is provided, the method comprising: performing methylation sequencing on a sample to be tested using methylation markers, primer pairs or probes disclosed herein; and determining, based on the methylation sequencing results, whether the sample to be tested has cancer or the type of cancer it has.
[0067] According to another aspect of this disclosure, a storage medium is provided in which the steps of the method for screening multiple cancer methylation markers, the apparatus for screening combinations of methylation markers, the multiple cancer detection model, or the multiple cancer detection system are implemented when the computer program is executed by a processor.
[0068] According to another aspect of this disclosure, a method for screening multiple cancer methylation biomarkers, an apparatus for screening combinations of methylation biomarkers, the methylation biomarkers, the primer pairs, the probes, the reagent kit, a method for constructing a model, the model, a method for detecting multiple cancers, a method for predicting organ origin, and the use of the multi-cancer detection system in the preparation of products for cancer detection are provided. In some embodiments, the use includes early cancer screening, diagnosis, auxiliary diagnosis, or prognosis prediction.
[0069] Beneficial effects:
[0070] This disclosure provides a comprehensive biomarker evaluation method that precisely identifies low-noise, high-sensitivity methylation biomarkers through a two-stage process. The selected methylation biomarkers are used to simultaneously detect combinations of methylation biomarkers for eight types of cancer: liver cancer, lung cancer, colorectal cancer, gastric cancer, esophageal cancer, pancreatic cancer, breast cancer, and ovarian cancer, and to precisely trace the tumor's origin. Specifically, it provides a method or system for assessing the cancer risk of a sample and predicting the tissue from which cancer signals originate in a patient.
[0071] This disclosure provides a comprehensive biomarker evaluation and screening method that precisely identifies low-noise, high-sensitivity methylation biomarkers through a two-stage process. The first stage evaluates biomarker sensitivity and fluid sample noise from both whole-genome methylation tissue sequencing and whole-genome methylation sequencing of healthy body fluid samples. This ensures a broad target range while reducing the inclusion of redundant biomarkers. In the initial screening stage, biomarkers are screened separately for each individual cancer type, effectively preventing the omission of biomarkers with good coverage and high sensitivity for a specific cancer type. The second stage further evaluates the discriminative performance of candidate biomarkers in body fluid samples through high-depth body fluid sample sequencing. Furthermore, in the early screening biomarker selection method, this invention proposes a supplementary biomarker method to enhance the complementarity between biomarkers and improve the detection rate of difficult-to-detect samples. This method effectively improves the detection sensitivity of cancer-related body fluid samples.
[0072] The liquid biopsy method disclosed herein uses DNA oligonucleotide probe sequences to target and capture the methylation regions involved in the combination, and predicts and assesses the presence and origin of tumor-derived ctDNA signals. It can significantly improve the sensitivity and specificity of multi-cancer diagnostic screening and achieve accurate source tracing. As a non-invasive detection technology, it is non-invasive, highly accurate, and has high patient compliance. Detecting ctDNA through liquid biopsy technology can detect cancer earlier and predict the likelihood of cancer occurrence. Predicting the source tissue of a patient's cancer signal has important auxiliary significance for determining downstream treatment pathways and saving medical costs. In summary, the liquid biopsy method disclosed herein has the potential to become an efficient early cancer screening tool, further reducing the medical burden and improving patient survival benefits. Attached Figure Description
[0073] Figure 1 A schematic diagram of the queue design in the embodiment is shown.
[0074] Figure 2 Details of the sensitivity of the validation set early screening model for each cancer type are shown in Example 4.
[0075] Figure 3 The Top1 prediction details for the source tracing of positive early screening samples in Example 4 are shown. Detailed Implementation
[0076] Hepatocellular carcinoma (HCC) is a common malignant tumor in my country, characterized by its insidious onset, high malignancy, and poor prognosis. Most HCC patients are diagnosed at an intermediate or advanced stage, and the recurrence rate is high. Currently, alpha-fetoprotein (AFP) combined with imaging examinations is commonly used for diagnosis. When AFP levels exceed the normal range, abdominal ultrasound or enhanced CT scans confirm the diagnosis. However, AFP has low sensitivity for early liver cancer screening and when combined with ultrasound examinations.
[0077] Lung cancer ranks first in both incidence and mortality among cancers in my country, seriously endangering people's lives and health. Most patients are diagnosed at an advanced stage, missing the optimal window for radical treatment. Therefore, early screening and diagnosis of lung cancer are crucial and can significantly improve patient survival rates. Clinically, imaging examinations are commonly used for lung cancer screening, including chest X-ray, chest LDCT, magnetic resonance imaging (MRI), and PET-CT, but their ability to screen for early-stage cancer remains insufficient. Colorectal cancer is a highly prevalent malignant tumor. Colonoscopy is the gold standard for colorectal cancer diagnosis, but it is painful, has poor patient compliance, and is not suitable for large-scale screening. In addition, the fecal immunochemical test (FIT) can be used, but its sensitivity is poor. Stomach and esophageal cancers are common malignant tumors of the digestive tract, seriously affecting patients' health and quality of life. Changes in dietary structure and poor eating habits have led to a year-on-year increase in the incidence of stomach and esophageal cancer. Due to the lack of obvious early symptoms, patients have a low rate of seeking medical attention, and by the time they are diagnosed, it is often at an advanced stage, missing the optimal treatment time, making treatment difficult and resulting in a poor prognosis. Currently, the most commonly used diagnostic methods in clinical practice are gastroscopy and esophagoscopy, which are difficult to use for large-scale screening to detect early-stage patients. Breast cancer and ovarian cancer are the most common malignant tumors in women, seriously endangering women's health. Ultrasound, mammography, and magnetic resonance imaging (MRI) are commonly used in clinical practice for breast cancer screening, but these methods are highly dependent on the doctor's skill level and have a high rate of missed diagnoses, making early screening challenging. Ovarian cancer can be screened clinically using ultrasound and MRI, but due to the lack of specific symptoms in the early stages, patient consultation rates are low, necessitating the search for simpler screening methods suitable for large populations. Pancreatic cancer is one of the most malignant tumors, with its incidence increasing year by year, a poor prognosis, and early-stage pancreatic cancer often asymptomatic and difficult to detect. Clinically, there is a lack of standardized screening methods, and diagnosis is mainly based on imaging examinations, which have significant limitations in early diagnosis. Therefore, for the screening and diagnosis of early cancers, it is necessary to find more effective detection methods that are more conducive to large-scale screening.
[0078] Currently, the main methods for cancer screening are imaging examinations and cancer marker detection, but these methods suffer from poor patient compliance, poor specificity and sensitivity, and lack of effective screening methods for some cancers.
[0079] Therefore, to further improve the detection rate of early screening for various cancers and accurately predict the site of cancer development, this disclosure provides a comprehensive biomarker screening method. First, whole-genome GMseq methylation sequencing is performed using cancer tissues and adjacent normal tissues associated with eight cancer types to initially screen for cancer-specific hypermethylated regions and tissue-specific hypermethylated regions. Further, to ensure lower noise levels in liquid biopsy and improve detection sensitivity, methylation noise is assessed using whole-genome GMseq methylation sequencing data from a baseline plasma cohort. Regions with high noise levels are eliminated, and probes are finally customized for the remaining regions.
[0080] Using the obtained region-customized probes, this disclosure performs high-depth targeted methylation detection on plasma samples related to eight cancer types and non-cancer plasma samples to further evaluate the performance of candidate biomarkers in plasma samples. When evaluating the plasma performance of biomarkers, it is necessary to ensure that the methylation signal in the aforementioned tissues is consistent with the methylation signal trend in plasma to reduce the risk of screening for false positive biomarkers. Furthermore, by evaluating the plasma performance of each candidate biomarker, the top-performing biomarkers are selected as the final biomarker combination.
[0081] This disclosure uses the above-described method to screen for methylation biomarker combinations related to early cancer screening and methylation biomarker combinations related to cancer tracing. In specific embodiments, the methylation biomarker combinations of this disclosure can be used for the identification and tracing of eight types of cancer: liver cancer, lung cancer, colorectal cancer, gastric cancer, esophageal cancer, pancreatic cancer, breast cancer, and ovarian cancer.
[0082] The methylation biomarker combinations involved in eight types of cancer screening and tracing include 466 early screening biomarker combinations in Table 1 and 1483 tracing biomarker combinations in Table 2. The physical locations of the early screening and tracing biomarkers were determined based on alignment with the human whole genome sequence (version hg19).
[0083] Table 1: Combinations of methylation biomarkers in early screening
[0084]
[0085]
[0086]
[0087]
[0088] Table 2: Combinations of Source-Tracing Methylation Biomarkers
[0089]
[0090]
[0091]
[0092]
[0093]
[0094]
[0095]
[0096]
[0097]
[0098]
[0099]
[0100] On the other hand, this disclosure also provides an early screening prediction model and a source tracing prediction model. First, this disclosure defines a positive read ratio index to assess the abundance of DNA reads from cancer tissue within the target biomarker region. Further, based on the panel high-depth methylation sequencing results of cancer blood and non-cancer blood, a threshold for the positive read ratio of each biomarker is defined, and the positive read ratio index is binary (1 if the positive read ratio of the target biomarker is higher than the corresponding threshold, and 0 otherwise).
[0101] In terms of building an early screening prediction model, the binary variable of the combination of early screening biomarkers provided in this disclosure is used as the input feature to build a logistic regression model (binary classification model). The LogisticRegression function in sklearn (V1.3.0) is used to build the model. Before training, a grid search method is used to traverse the hyperparameter combination to lock the optimal hyperparameters. Furthermore, a binary classification model is built and the early screening positive threshold is determined by the 5-fold cross-validation method.
[0102] In terms of constructing the source tracing prediction model, the binary variables of the source marker combination provided in this disclosure are used as input features to construct a logistic regression model (multi-class model). The LogisticRegression function in sklearn (V1.3.0) is used for model construction. Before training, a grid search method is used to traverse the hyperparameter combination to lock the optimal hyperparameters. Furthermore, a multi-class model is constructed using the 5-fold cross-validation method. The organ with the highest probability is selected as the Top1 predicted organ (the organ with the most likely source), and the organ with the second highest probability is selected as the Top2 predicted organ (the organ with the second most likely source).
[0103] Finally, this disclosure includes additional plasma samples from 31 liver cancer patients, 29 lung cancer patients, 42 colorectal cancer patients, 15 gastric cancer patients, 36 esophageal cancer patients, 28 pancreatic cancer patients, 24 ovarian cancer patients, 20 breast cancer patients, 148 healthy plasma samples, and 43 benign plasma samples as an independent validation set to verify the performance of the early screening prediction model and the source tracing prediction model when using the above-mentioned biomarker combinations.
[0104] The constructed multi-cancer early screening model and source tracing model, when used in combination with the above biomarkers, can simultaneously and accurately detect eight types of cancer signals and perform precise source tracing. While ensuring a low false positive rate, it significantly improves the detection rate of these eight types of cancer patients and accurately predicts the source organ of cancer signals.
[0105] definition
[0106] Unless otherwise defined, all technical and scientific terms used in this disclosure have the same meaning as commonly used in the field to which this disclosure pertains. For purposes of interpreting this specification, the following definitions will apply, and where appropriate, terms used in the singular will also include the plural form, and vice versa.
[0107] Unless the context clearly indicates otherwise, the terms “a” and “an” as used herein include plural references.
[0108] The term "about" as used herein is as understood by one of ordinary skill in the art and varies within a certain range depending on the context in which it is used. If one of ordinary skill in the art is unfamiliar with the use of this term in the context in which it is used, "about" will mean a particular value plus or minus 10%.
[0109] As used herein, the term "cfDNA molecule" or "cfDNA" includes DNA molecules that are naturally present in a subject in an extracellular form (e.g., in blood, serum, plasma, or other bodily fluids such as lymph, cerebrospinal fluid, urine, or sputum). Although cfDNA is originally present in one or more cells of a large, complex biological organism (e.g., a mammal), cfDNA undergoes release from the cell into the fluid present in the organism, and can therefore be obtained by obtaining a sample of the fluid without the need for an in vitro cell lysis step.
[0110] The term "CpG site" used in this article is an abbreviation for "-C-phosphate-G-", also known as a "CG site," referring to a region on the DNA sequence where the base sequence is cytosine followed by guanine, linked by a phosphodiester bond. C is located at the 5' end and G at the 3' end. Cytosine at CpG sites can be methylated to 5-methylcytosine. In mammals, 70% to 80% of CpG sites have methylated cytosine.
[0111] As used herein, the term "methylation" or "DNA methylation" refers to the addition of a methyl group to a nucleotide base in a nucleic acid molecule. In some embodiments, methylation involves adding a methyl group to cytosine at a CpG site (cytosine-phosphate-guanine site, i.e., cytosine followed by guanine in the 5'-3' direction of the nucleic acid sequence). In some embodiments, DNA methylation involves adding a methyl group to adenine, such as N6-methyladenine. In some embodiments, DNA methylation is 5-methylation (modification of the 5th carbon of cytosine). In some embodiments, 5-methylation involves adding a methyl group to the 5C position of cytosine to generate 5-methylcytosine (5mC). In some embodiments, methylation includes derivatives of 5mC. Derivatives of 5mC include, but are not limited to, 5-hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), and 5-carboxycytosine (5-caC). In some embodiments, DNA methylation is 3C methylation (modification of the 3rd carbon of cytosine). In some implementations, 3C methylation involves adding a methyl group to the 3C position of cytosine to generate 3-methylcytosine (3mC). Methylation can also occur at non-CpG sites; for example, methylation can occur at CpA, CpT, or CpC sites. DNA methylation can alter the activity of methylated DNA regions. For example, when DNA in a promoter region is methylated, gene transcription may be repressed. DNA methylation is essential for normal development, and abnormalities in methylation can disrupt epigenetic regulation. Disruptions in epigenetic regulation, such as repression, can cause diseases such as cancer. Promoter methylation in DNA can indicate cancer.
[0112] The term “deep sequencing” as used in this article refers to the general concept of a large number of repeated reads for each region of a sequence.
[0113] As used herein, the term "sequencing data" refers to any sequence information about a nucleic acid molecule known to a person skilled in the art. Sequencing data may include information about the DNA or RNA sequence that must be converted into a nucleic acid sequence, modified nucleic acids, single-stranded or double-stranded sequences, or alternatively, amino acid sequences. Sequencing data may additionally include information about the sequencing equipment, acquisition date, read length, sequencing orientation, source of the sequenced entity, adjacent sequences or reads, the presence of duplicates, or any other suitable parameters known to a person skilled in the art. Sequencing data may be presented in any suitable format, file, encoding, or document known to a person skilled in the art.
[0114] As used herein, the term "tumor" refers to a mass or growth that is defined by itself as an abnormal new growth of cells, which typically grow faster than normal cells and will continue to grow if left untreated, sometimes causing damage to adjacent structures. Tumors can vary greatly in size. Tumors can be solid or fluid-filled. Tumors can refer to benign (non-malignant, usually harmless) or malignant (capable of metastasis) growths. Some tumors may contain benign neoplastic cells (e.g., carcinoma in situ) as well as malignant cancer cells (e.g., adenocarcinoma). It should be understood that this includes growths located in multiple locations throughout the body. Therefore, for the purposes of this disclosure, tumors include primary tumors, lymph nodes, lymphatic tissue, and metastatic tumors.
[0115] Non-limiting examples of the cancers mentioned include bile duct cancer, bladder cancer, transitional cell carcinoma, urothelial carcinoma, breast cancer, cervical cancer, cervical squamous cell carcinoma, rectal cancer, colorectal cancer, colon cancer, hereditary nonpolyposis colorectal cancer, colorectal adenocarcinoma, gastrointestinal stromal tumor (GIST), endometrial cancer, endometrial stromal sarcoma, esophageal cancer, esophageal squamous cell carcinoma, esophageal adenocarcinoma, ocular melanoma, uveal melanoma, gallbladder cancer, gallbladder adenocarcinoma, renal cell carcinoma, and clear cell renal cell carcinoma. Transitional cell carcinoma, urothelial carcinoma, nephroblastoma, liver cancer, hepatocellular carcinoma, cholangiocarcinoma, hepatoblastoma, lung cancer, non-small cell lung cancer (NSCLC), nasopharyngeal carcinoma (NPC), neuroblastoma, oral cancer, oral squamous cell carcinoma, ovarian cancer, pancreatic cancer, pancreatic ductal adenocarcinoma, pseudopapillary tumor, acinar cell carcinoma, prostate cancer, skin cancer, melanoma, malignant melanoma, cutaneous melanoma, small intestine cancer, gastric cancer, gastric epithelial carcinoma, uterine cancer or uterine sarcoma.
[0116] As used in this article, the term “sequencing” refers to the process of determining the sequence (e.g., the identity and order of monomeric units) of a biomolecule, such as a nucleic acid, like DNA or RNA. Exemplary sequencing methods include, but are not limited to, targeted sequencing, single-molecule real-time sequencing, exon or exome sequencing, intron sequencing, electron microscopy-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxy termination sequencing, whole genome sequencing, hybridization sequencing, pyrosequencing, capillary electrophoresis, double-strand sequencing, cyclic sequencing, single-base extension sequencing, solid-phase sequencing, high-throughput sequencing, massively parallel signature sequencing, emulsion PCR, low denaturing temperature co-amplification PCR (COLD-PCR), multiplex PCR, reversible dye terminator sequencing, paired-end sequencing, near-term sequencing, exonuclease sequencing, ligation sequencing, short-read sequencing, single-molecule sequencing, synthetic sequencing, real-time sequencing, reverse terminator sequencing, nanopore sequencing, 454 sequencing, Solexa genome analyzer sequencing, SOLiD™ sequencing, MS-PET sequencing, DNA nanosphere sequencing (DNBSEQ), complex probe anchored polymerization sequencing (cPAS), and combinations thereof. In some implementations, sequencing can be performed using a gene analyzer, such as those commercially available from Illumina, Inc., Pacific Biosciences, Inc., Applied Biosystems / Thermo Fisher Scientific, or BGI Genomics Co., Ltd. Examples include BGI's DNBseq sequencing platforms such as BGISEQ-500, BGISEQ-50, MGISEQ-2000, MGISEQ-200, DNBSEQ-T7, DNBSEQ-G99, and DNBSEQ-T20X2, or Illumina's HiSeq2000, HiSeq2500, HiSeq4000, HiSeqX10, and NovaSeq6000.
[0117] Sequencing can be performed by any method known in the art. DNA sequencing techniques include classical dideoxy sequencing (Sanger method) using labeled terminators or primers, synthetic sequencing using labeled nucleotides with reversible termination, pyrosequencing, 454 sequencing, allele-specific hybridization with labeled oligonucleotide probe libraries, synthetic sequencing using allele-specific hybridization with labeled clonal libraries followed by ligation, real-time monitoring of labeled nucleotide incorporation during the polymerization step, polymerase cloning sequencing, and SOLiD sequencing. Sequencing of isolated molecules has recently been demonstrated using sequential or single extension reactions with polymerases or ligases, and by single or sequence differential hybridization with probe libraries.
[0118] The term “deep sequencing” as used in this article refers to the general concept of a large number of repeated reads for each region of a sequence.
[0119] As used herein, the term "sequencing data" refers to any sequence information about a nucleic acid molecule known to a person skilled in the art. Sequencing data may include information about the DNA or RNA sequence that must be converted into a nucleic acid sequence, modified nucleic acids, single-stranded or double-stranded sequences, or alternatively, amino acid sequences. Sequencing data may additionally include information about the sequencing equipment, acquisition date, read length, sequencing orientation, source of the sequenced entity, adjacent sequences or reads, the presence of duplicates, or any other suitable parameters known to a person skilled in the art. Sequencing data may be presented in any suitable format, file, encoding, or document known to a person skilled in the art.
[0120] As used in this article, the terms "sequencing read," "read," or "segment" refer to a sequencing read derived from a nucleic acid sample. Typically, although not strictly necessary, a read represents a sequence of consecutive base pairs in the sample. A read can be symbolically represented by the sequence of base pairs in the A, T, C, and G portions of the sample, along with a probability estimate of the correctness of those base pairs (quality score). It can be stored in a memory device and processed as needed to determine if it matches a reference sequence or meets other criteria. Reads can be obtained directly from sequencing equipment or indirectly from stored sequence information about the sample.
[0121] As used herein, the term "subject" refers to any animal, mammal, or human. The subject has, may have, or is suspected of having any one or more combined diseases. The subject may have cancer, may exhibit cancer-related symptoms, may not exhibit cancer-related symptoms, or may not have been diagnosed with cancer. In some embodiments, the subject is a human being.
[0122] As used herein, the terms "computer-readable medium" (e.g., data storage, data storage, etc.) or "computer-readable storage medium" refer to any medium that participates in providing instructions to a processor for execution. Such media can take many forms, including but not limited to non-volatile media, volatile media, and transmission media. Examples of non-volatile media include, but are not limited to, optical discs, solid-state drives, and magnetic disks, such as storage devices. Examples of volatile media include, but are not limited to, dynamic memory, such as RAM.
[0123] Common forms of computer-readable media include, for example, floppy disks, floppy disks, hard disks, magnetic tapes or any other magnetic media, CD-ROMs, any other optical media, punched cards, paper tapes, any other physical media with a perforated pattern, RAM, PROMs and EPROMs, FLASH-EPROMs, any other memory chips or cassette tapes, or any other tangible media from which a computer can read.
[0124] In addition to computer-readable media, data may be provided as signals on a transmission medium included in a communication device or system to provide one or more sequences of instructions to a processor of a computer system for execution. For example, a communication device may include a transceiver having signals indicating instructions and data. The instructions and data are configured to cause one or more processors to perform the functions outlined in this disclosure. Representative examples of data communication transmission connections may include, for example, telephone modem connections, wide area networks (WANs), local area networks (LANs), infrared data connections, NFC connections, etc.
[0125] The term "diagnosis" as used herein refers not only to "judgment" but also includes "determination," "differentiation," "identification," "examination," "detection," and "prediction." This includes differentiating or assisting in differentiating / identifying cancer and benign lesions, determining / assessing cancer and its progression, and predicting cancer risk, treatment effectiveness, prognosis, and recurrence risk. In some embodiments, the diagnosis includes differentiating or assisting in differentiating cancer and benign lesions, determining the progression of cancer, or predicting cancer risk, treatment effectiveness, prognosis, and recurrence risk.
[0126] The following embodiments and accompanying drawings are provided to aid in understanding this disclosure. However, it should be understood that these embodiments and drawings are for illustrative purposes only and do not constitute any limitation. The actual scope of protection of this disclosure is set forth in the claims. It should be noted that various modifications and improvements made by those skilled in the art based on this inventive concept are within the scope of protection of this disclosure. In the following description, descriptions of well-known structures and techniques are omitted to avoid unnecessarily obscuring the concepts of this disclosure. Such structures and techniques have also been described in many publications. Furthermore, reagents used, unless otherwise specified, are all commercially available conventional products.
[0127] Example
[0128] The queue design diagrams in Examples 1-5 are as follows: Figure 1 As shown.
[0129] The methylation sequencing involved in the examples includes the following steps:
[0130] 1. Sample DNA extraction
[0131] 1.1 Plasma Sample Extraction
[0132] For whole blood, timely plasma / blood cell separation should be performed (EDTA anticoagulant tubes, within 4 hours; Streck tubes, within 72 hours). The separation steps are as follows:
[0133] ① Centrifuge at 1600g for 10 minutes at 4℃. After centrifugation, aliquot the upper plasma into multiple 1.5mL or 2.0mL centrifuge tubes. When aspirating the plasma, be careful not to aspirate the white blood cells in the middle layer.
[0134] ② Centrifuge at 16000g for 10 minutes at 4℃ to remove residual cells, and transfer the supernatant into a new 1.5mL or 2.0mL centrifuge tube (be careful not to aspirate the white blood cells at the bottom of the tube) to obtain the required plasma.
[0135] plasma according to MagMAX TM Follow the instructions for the Cell-Free DNA Isolation Kit to extract cfDNA from plasma.
[0136] 1.2 Tissue Sample Extraction
[0137] DNA was extracted from cancerous tissue and paired adjacent normal tissue using the RNeasy Mini Kit (Qiagen). 100 ng gDNA was then sonicated to break it down, and the broken DNA was retained for subsequent library construction and sequencing experiments.
[0138] 2. Construction of methylated libraries
[0139] Using Hieff The Ultima Pro DNA Library Prep Kit for Illumina is used to construct a library from extracted DNA, followed by TET enzyme oxidation and pyridine boronane reduction, and then PCR amplification to build a pretext library.
[0140] 3. Hybridization and Sequencing
[0141] Whole-genome methylation sequencing does not require hybridization; however, targeted methylation sequencing requires prior hybridization capture of the library before sequencing using the Gene+seq sequencer or other sequencers based on the same principle. Sequencing procedures should be performed according to the manufacturer's instructions.
[0142] Example 1: Screening of cancer biomarker combinations
[0143] The final biomarker combination was obtained by screening methylation sequencing data from two separate stages.
[0144] In the first phase, 733 tissue samples and 271 healthy plasma samples as described in Table 3 were used for 30× and 60× GMseq whole genome methylation sequencing, respectively, to preliminarily determine the candidate biomarker set and to customize probes according to the target regions.
[0145] Table 3. Information on tissue samples detected in Example 1
[0146] Cancer Number of cancer tissue samples Number of adjacent normal tissue samples liver cancer 50 12 Colorectal cancer 37 23 lung cancer 97 210 Ovarian cancer 119 20 pancreatic cancer 36 18 Breast cancer 35 5 Stomach cancer 24 7 esophageal cancer 32 8
[0147] In the second phase, the customized probes from the first phase were used to further evaluate the distinguishing performance of the biomarkers in plasma using targeted methylation sequencing results from 1405 plasma samples in the training set described in Table 4, and to screen out the final biomarker combination. This sequencing result will be applied to both phases; the first part of the data (training set) shown in Table 4 was used for biomarker plasma screening and model training, including 530 cancer plasma samples and 459 non-cancer plasma samples (details in Table 4); the second part of the data (validation set) shown in Table 4 was used to independently validate the clinical predictive performance of the early screening model and the source tracing model, including 225 cancer plasma samples and 191 non-cancer plasma samples.
[0148] Table 4. Clinical information of plasma samples detected in Example 1
[0149]
[0150] The specific steps of the above screening method are as follows:
[0151] 1. Identification of hypermethylated regions related to early tumor screening and source tracing (Phase 1)
[0152] First, before screening for highly methylated regions, this disclosure uses wgbstools software. The software analysis parameters are: the reference genome is hg19, the methylation analysis results of the cancer tissue are input, and the human genome is divided into 2.02 million regions based on the whole-genome methylation sequencing results of the cancer tissue in Table 3. Each pre-divided region is used as a marker region for subsequent screening analysis.
[0153] Second, this disclosure employs a cancer-by-cancer screening and merging strategy in determining the hypermethylated regions related to early screening and source tracing. Specifically, it screens for cancer-related specific hypermethylated regions and tissue-related specific hypermethylated regions for each cancer. First, for each cancer type, matching cancer tissues and adjacent normal tissues are selected, and a set of early screening biomarkers, A, is determined through regional difference analysis (e.g., average methylation rate difference greater than 0.2 and P-value less than 0.01). Second, for each cancer type, cancer tissues of the target cancer type and other cancer-related tissues (including cancer tissues and adjacent normal tissues) are selected, and a "one-to-one" strategy is used n times to screen for cancer tissue-specific hypermethylated regions of the target cancer type (i.e., 7 difference analyses are performed, with the background group for each difference analysis being cancer tissues and adjacent normal tissues of other cancer types; for example, each difference analysis result is required to meet the requirement that the average methylation rate difference is greater than 0.2 and P-value less than 0.01), which serve as the set of source tracing biomarkers, B. The P-value was calculated using the Wilcoxon rank-sum test; the average methylation rate was the proportion of the number of methylated CpG sites on all reads within the target marker region to the total number of detectable CpG sites.
[0154] Third, high-noise regions were further filtered from the whole-genome methylation sequencing results of 271 healthy plasma samples. First, the regional mean methylation rate distribution of markers in sets A and B in the 271 healthy plasma samples was evaluated, requiring that the regional mean methylation rate of the 95th percentile be less than 0.2. The markers filtered under this condition were designated as sets C (from set A) and D (from set B).
[0155] Fourth, probes are customized based on the regions of biomarkers in sets C and D for subsequent targeted methylation high-depth sequencing and further determination of early screening plasma biomarker combinations and traceable plasma biomarker combinations.
[0156] 2. Determination of the combination of early cancer screening biomarkers (Phase II)
[0157] First, the early screening biomarker combination included in this disclosure follows a cancer-by-cancer screening and merging strategy. Biomarkers are screened separately for each cancer type (a total of eight cancer types). Finally, the biomarkers screened for the eight cancer types are summarized and merged, and duplicates are removed to form the final biomarker combination. The screening method for biomarkers for each cancer type is consistent. Therefore, this embodiment only uses colorectal cancer as an example. The screening data in this stage all use the targeted methylation high-depth sequencing results of plasma samples in the training set in Table 4 (the target region is the probe region determined in the first stage, i.e., the screening set comes from set C). "High-depth sequencing" refers to the original sequencing depth of 6000× or more. The original depth refers to the total amount of sequence data generated by sequencing in the probe region divided by the probe size.
[0158] Second, taking colorectal cancer as an example, this embodiment uses 80 cases of colorectal cancer plasma as positive samples and 459 cases of non-cancer plasma as negative samples. Further, by statistically analyzing the extreme hypermethylation read ratios of each biomarker in the 459 non-cancer plasma cohort, and using the 95th percentile of the extreme hypermethylation read ratio as the matching threshold for each biomarker, if the extreme hypermethylation read ratio of a biomarker in the target sample is higher than the corresponding threshold, it is recorded as 1; otherwise, it is recorded as 0. This achieves binary representation of each biomarker variable, allowing for further evaluation based on the binary values of the biomarkers. Estimate the coverage of each biomarker in the target cancer blood sample (a biomarker binary value of 1 indicates coverage of the cancer blood sample, and 0 indicates no coverage); where "extremely hypermethylated read ratio" is the proportion of extremely hypermethylated reads in the genomic region covered by the biomarker to the total reads. "Extremely hypermethylated reads" are reads in the target region where the number of methylated CpG sites is not less than 0.8 of the total number of CpG sites and covers at least 3 CpGs. "Total reads" are reads in the target region that cover at least 3 CpGs.
[0159] Third, the coverage of biomarkers in the colorectal cancer blood cohort was evaluated using 80 colorectal cancer blood samples. The coverage rate (the proportion of blood samples with a binary value of 1 for the biomarker to the total number of blood samples) was calculated by dividing the coverage rate by the ratio of the coverage rate to the 80 colorectal cancer blood samples. The top 128 biomarkers with coverage rates from high to low and a coverage rate greater than 5% were selected as the screening results.
[0160] Fourth, the screening results of eight cancer early screening biomarkers were merged and deduplicated to obtain set E1. Further, based on the characteristics of tumor heterogeneity, a self-developed iterative algorithm was used to supplement some biomarkers from set C to improve the overall detection sensitivity of the biomarker combination (this part uses all cancer blood samples and all non-cancer blood samples from the training set). The self-developed iterative algorithm first defines the complementary gain score Qi for each new biomarker. Qi is determined by the positive sample distribution of biomarker i (if biomarker i is higher than the matching threshold in the plasma sample, it is considered to cover the positive sample). The plasma samples in the training set are divided into three levels and corresponding to three weight coefficients based on the biomarker combination before each iteration. The proportion of positive biomarkers in each plasma sample before the addition of biomarkers is calculated (i.e., the proportion of biomarkers higher than the matching threshold in the plasma sample). The 60th percentile of the proportion of positive biomarkers in cancer plasma samples is used as the threshold P1. The weight of samples with a proportion of positive biomarkers lower than P1 in cancer blood samples is set as follows: 1. Using the 90th percentile of the proportion of positive biomarkers in non-cancer plasma samples as threshold P2 and the 95th percentile of the proportion of positive biomarkers in non-cancer plasma samples as threshold P3, the weight of samples with a positive biomarker proportion higher than P3 in non-cancer plasma samples is set to -2, and the weight of samples with a positive biomarker proportion higher than P2 but lower than P3 in non-cancer plasma samples is set to -1. The three sample weight coefficients change dynamically with the change of biomarker combination. When calculating the Qi score, only plasma samples that meet the above three conditions are included. The Qi score is used to dynamically evaluate the complementary gain score of the newly added biomarker i under the current situation. The higher the value, the greater the sensitivity performance improvement that the newly added biomarker can bring.
[0161]
[0162] Where the weight coefficient W1 is 1, the weight coefficient W2 is -1, the weight coefficient W3 is -2, and N C N represents cancer blood samples in the original biomarker combination whose positive biomarker percentage is lower than P1. h1 N represents the proportion of non-cancer samples with a positive biomarker higher than P2 but lower than P3 in the original biomarker combination cohort. h2 R represents the proportion of non-cancer samples with a higher percentage of positive biomarkers than P3 in the original biomarker combination cohort. in The proportion of extremely high methylation reads representing the new biomarker i in n samples, C i The matching threshold represents the ratio of extremely hypermethylated reads of biomarker i, wherein the positive biomarker is a biomarker with an extremely hypermethylated read ratio higher than the matching threshold, and the matching threshold is the 95th percentile of the proportion of extremely hypermethylated fragments of each biomarker in the aforementioned 459 non-cancer plasma samples.
[0163] Score Q using complementary gain iIn each iteration, the marker with the highest score can be selected as the new marker, and two conditions can be set as iteration termination conditions: ①Q i When the value is less than or equal to 0, the iteration process of adding new biomarkers is stopped. ② Using the 99th percentile of the proportion of positive biomarkers in non-cancer samples as the threshold, when the number of cancer blood samples above the threshold after adding a biomarker is less than the number of cancer blood samples above the threshold before adding a biomarker, the iteration process of adding new biomarkers is stopped.
[0164] Fifth, the final early screening marker combination is the combination of the merged marker set E1 and the supplementary marker set E2 selected by the above iterative process, which contains a total of 466 methylation regions. The detailed results are shown in Table 1.
[0165] 3. Determination of the combination of tumor biomarkers (Phase II)
[0166] First, the source biomarker combinations included in this publication follow a cancer-specific screening and merging strategy. The difference from the early screening biomarker screening method is that the control group is changed from non-cancer plasma samples to blood samples from other cancers. Biomarkers are screened separately for each cancer type and then finally summarized and merged as the final source biomarker combination. The source biomarker combination screening method is consistent for each cancer type, so colorectal cancer is used as an example. The screening data comes from the targeted methylation high-depth sequencing results of clinical plasma samples in the training set, and the screening scope comes from set D.
[0167] Second, taking colorectal cancer as an example, this embodiment uses plasma from 80 cases of colorectal cancer as positive samples, and plasma from 77 cases of lung cancer, 71 cases of liver cancer, 60 cases of pancreatic cancer, 71 cases of esophageal cancer, 59 cases of gastric cancer, 56 cases of breast cancer, and 64 cases of ovarian cancer as negative sample cohorts. Further, a strategy of analyzing each pair of combinations is employed: colorectal cancer vs. lung cancer, colorectal cancer vs. liver cancer, colorectal cancer vs. pancreatic cancer, colorectal cancer vs. esophageal cancer, colorectal cancer vs. gastric cancer, colorectal cancer vs. breast cancer, and colorectal cancer vs. ovarian cancer. Under each pair of combinations, the extreme hypermethylation read ratio of each marker in the negative cancer blood cohort is statistically analyzed, and the 96th percentile of the extreme hypermethylation read ratio is used as the threshold for each marker under that combination. If the extreme hypermethylation read ratio of a marker in the target sample is higher than this threshold, it is recorded as 1; otherwise, it is recorded as 0. This achieves the binary representation of each marker variable.
[0168] Third, for each pair of samples, biomarkers associated with tissue-specific expression in colorectal cancer are evaluated. The selectable biomarkers are derived from the set of biomarkers that are specifically highly expressed in colorectal cancer tissue in set D (to ensure that the selected biomarkers have consistent high methylation signal characteristics in plasma and tissue). The top 64 biomarkers with the highest coverage and the absolute positive coverage number greater than 5 are selected as the screening results. The "absolute positive coverage number" is the coverage number of the biomarker in a body fluid sample of a certain cancer type minus the coverage number of the biomarker in any other body fluid sample of a different cancer type.
[0169] Fourth, the results of screening for each cancer combination were merged and deduplicated, retaining a total of 1483 methylation regions. The details are shown in Table 2.
[0170] Example 2: Construction of an Eight-Cancer Early Screening Prediction Model
[0171] The main purpose of this embodiment is to construct an early screening model using the combination of early screening cancer biomarkers proposed in Example 1. A multi-cancer early screening model was constructed using 989 clinical plasma samples that underwent targeted methylation detection in the training set of Example 1. A total of 530 cancer plasma samples (including 69 lung cancer, 71 liver cancer, 80 colorectal cancer, 60 pancreatic cancer, 71 esophageal cancer, 59 gastric cancer, 56 breast cancer, and 64 ovarian cancer) and 459 non-cancer plasma samples were included.
[0172] First, for the early screening biomarker combination in Example 1, a metric for measuring the abnormal methylation signal of each biomarker is proposed: positive read ratio. The positive read ratio for each biomarker region is the ratio of reads with positive methylation status in this region to all detected reads. The criteria for judging the positive methylation status of a read are as follows: ① The read covers at least 3 CpG sites and the methylation rate of the read is not less than 0.8; ② The probability of the methylation pattern of the read appearing in the baseline methylation pattern database is less than 0.1%. The methylation rate of the read is the proportion of the number of methylated CpG sites on the read to the total number of detectable CpG sites on the read.
[0173] The baseline methylation pattern database was constructed using targeted methylation sequencing data from 103 non-cancer plasma samples via a Markov model.
[0174] Second, the positive read ratio of methylation markers (the early screening marker combination proposed in Example 1) for each plasma sample (including cancer plasma samples and non-cancer plasma samples) was obtained using the above method. The 95th percentile of the positive read ratio results of 459 non-cancer plasma samples was then used as the positive threshold of the early screening marker combination (the early screening marker combination proposed in Example 1). The positive read ratio of each marker for all plasma samples (including cancer plasma samples and non-cancer plasma samples) could be analyzed using the positive threshold of each marker. Thus, the positive read ratio index of the early screening markers of all samples was transformed into a binary variable (1 if it is higher than the matching threshold, and 0 if it is lower than or equal to the matching threshold).
[0175] Third, the early screening markers of all samples were combined and transformed into a binary variable feature matrix, which was then used as input features to further construct a binary classification machine learning model. The LogisticRegression function in sklearn (V1.3.0) was used for model construction. The model parameters were continuously optimized and finally locked using grid search and cross-validation. The locked model was then used as the early screening model, and the 99.1% quantile of the prediction results of 459 non-cancer plasma samples was used as the threshold for cancer prediction. If the model result was higher than the threshold, it was predicted as cancer; if it was lower than the threshold, it was predicted as no cancer. The threshold of the model was 0.666.
[0176] Fourth, the predictive performance of the early screening model on the training set was initially evaluated through cross-validation. With a threshold of 0.666, the specificity was 99.1% and the sensitivity was 71.9%.
[0177] Example 3: Construction of the Source Tracing Prediction Model
[0178] The main purpose of this embodiment is to construct a traceability model using the traceability biomarker combination proposed in Embodiment 1. The traceability model is constructed using 530 clinical cancer plasma samples that underwent targeted methylation detection in Embodiment 1, including 69 cases of lung cancer, 71 cases of liver cancer, 80 cases of colorectal cancer, 60 cases of pancreatic cancer, 71 cases of esophageal cancer, 59 cases of gastric cancer, 56 cases of breast cancer, and 64 cases of ovarian cancer.
[0179] First, similar to the method in Example 2, the positive read ratio of each clinical plasma sample in 530 clinical cancer plasma samples was obtained on the combination of methylation markers (the source marker combination proposed in Example 1). The 95th percentile of the positive read ratio of the negative cancer blood group (the other 7 types of cancer blood samples that are not the target cancer) was used as the positive threshold for each marker. By using the positive threshold of each marker, the positive read ratio of each source marker of all samples can be converted into a binary variable (1 if it is higher than the threshold, and 0 if it is lower than or equal to the threshold).
[0180] Second, by combining the source markers of all samples into a binary variable matrix, a multi-class machine learning model is further constructed. Using the LogisticRegression function in sklearn (V1.3.0), a one-vs-rest strategy is adopted to construct the multi-class model. The model parameters are continuously optimized and finally locked by grid search and cross-validation. Finally, the locked multi-class model is used as the source model. The organ category with the highest predicted probability is selected as the Top1 source prediction result, and the organ category with the second highest probability is selected as the Top2 source prediction result.
[0181] Third, the prediction performance of the training set source tracing model was initially evaluated through cross-validation. The overall Top 1 score was 82%, and the overall Top 2 score was 90%. Only the tumor samples predicted as positive in Example 2 were used in the statistical analysis.
[0182] Example 4: Performance Verification of Early Screening Prediction Model and Source Tracing Prediction Model
[0183] The main purpose of this embodiment is to validate the performance of the early screening and source tracing model using 225 cancer blood samples (29 lung cancer, 31 liver cancer, 42 colorectal cancer, 28 pancreatic cancer, 36 esophageal cancer, 15 gastric cancer, 20 breast cancer, and 24 ovarian cancer) and 191 non-cancer plasma samples, which were not used in Embodiments 2 and 3 (validation set) in Embodiment 1.
[0184] First, the early screening prediction model from Example 2 and the source tracing prediction model from Example 3 were used to perform early screening prediction and source tracing prediction on samples in the validation set, respectively. Statistical prediction results showed that the specificity of the early screening prediction model in the independent validation set was 99.0%, and the sensitivity was 74.7%. Details of the sensitivity for each cancer type and stage are as follows... Figure 2 As shown, the Top1 source tracing accuracy for early screening positive tumor samples was 84%, and the Top2 source tracing accuracy was 92%. Details of Top1 prediction for each cancer type are as follows: Figure 3 As shown in Table 5, the sensitivity of the early screening model and the overall performance of the source tracing model are compared.
[0185] Table 5: Sensitivity of Early Screening Model and Overall Performance of Source Tracing Model
[0186]
[0187]
[0188] Example 5: Comparison of the sensitivity of early screening markers before and after supplementing markers using the iterative algorithm
[0189] To fully demonstrate the improvement effect of the iterative algorithm on blood samples with difficult-to-detect cancer (stage I), the early screening model was trained using the iterative algorithm to supplement the biomarker combinations (E1 and E1+E2) before and after supplementation, respectively, in 530 cancer plasma samples (including 69 lung cancer, 71 liver cancer, 80 colorectal cancer, 60 pancreatic cancer, 71 esophageal cancer, 59 gastric cancer, 56 breast cancer, and 64 ovarian cancer) and 459 non-cancer plasma samples. To ensure comparability, the model thresholds were determined by the 99.1% quantile of the prediction results of the 459 non-cancer plasma samples. Furthermore, the performance changes of the early screening biomarker combination before and after optimization were evaluated using 225 cancer blood samples (29 lung cancer, 31 liver cancer, 42 colorectal cancer, 28 pancreatic cancer, 36 esophageal cancer, 15 gastric cancer, 20 breast cancer, and 24 ovarian cancer) and 191 non-cancer plasma samples from Example 4 as independent validation cohorts. Based on the prediction of the above model and threshold, the validation set results are shown in Table 6. The specificity was 99.0% before and after supplementation, but the sensitivity was significantly improved after supplementation compared with that before supplementation.
[0190] Table 6: Sensitivity of Early Screening Models Before and After Supplementation with Biomarkers
[0191] plasma types Sensitivity before supplementation Sensitivity after supplementation Increase percentage All phases 72% 74.7% 3.8% Phase I 44.6% 53.6% 20.2% Phase II 69.0% 69.0% 0% Phase III 82.2% 83.6% 1.7% Phase IV 90.7% 90.7% 0%
[0192] The percentage increase is calculated as follows: (Sensitivity after supplementation - Sensitivity before supplementation) × 100% / Sensitivity before supplementation.
[0193] The technical solutions disclosed herein are not limited to the specific embodiments described above. Any technical modifications made based on the technical solutions disclosed herein shall fall within the protection scope of this disclosure.
Claims
1. A method for screening methylation markers, the method comprising: S1) Obtain the methylation sequencing results of cancer tissue samples and non-cancer tissue samples corresponding to the cancer type, and divide the genome into multiple methylation regions based on the methylation sequencing results; S2) Perform differential analysis on the methylation regions of cancer tissue samples and non-cancer tissue samples of the cancer type, screen out the methylation regions whose average methylation rate difference of the cancer type is greater than a threshold fraction relative to the non-cancer tissue samples, and merge and remove duplicates to obtain the biomarker set A. The methylation regions with average methylation rate differences greater than the threshold fraction for the cancer types include those with significantly high methylation rates. Step S2) further includes the step of filtering the set of markers A; The filtering step includes: selecting a set A' of markers from the marker set A that has an average methylation rate of less than 0.5 in non-cancer body fluid samples; S3) By performing differential analysis on the methylation regions of cancer tissue samples of the stated cancer type and cancer tissue samples of one or more other cancer types, methylation regions with an average methylation rate difference greater than a threshold fraction relative to cancer tissue samples of one or more other cancer types are screened out, and after merging and deduplication, a set of markers B is obtained. Step S3) further includes a step of filtering the biomarker set B; the filtering step includes: screening out biomarker set B' from biomarker set B that has an average methylation rate of less than 0.5 in tissue samples of one or more other cancer types; S4) Targeted methylation sequencing is performed on the biomarker set A' in the body fluid sample and non-cancer body fluid sample of the cancer type to obtain the coverage E1 of each biomarker in the cancer type, and the biomarker set C is obtained based on the coverage E1. The process of obtaining the coverage E1 of each biomarker in the cancer type includes: selecting reads with a methylation rate of 0.5 or higher and covering more than 3 CpG sites as extremely hypermethylated reads; the proportion of an extremely hypermethylated read of a biomarker to the total reads of that biomarker as the extreme hypermethylated read ratio of that biomarker; and binarying the extreme hypermethylated read ratios of each biomarker in the obtained cancer type body fluid samples based on threshold quantiles to determine the coverage E1 of each biomarker. The determination of the threshold quantiles includes: statistically analyzing the extreme hypermethylation read ratios of each marker in non-cancer blood samples, and using the 90% to 99% quantiles of the extreme hypermethylation read ratios as the threshold quantiles of each marker; The marker set C obtained based on coverage E1 includes: selecting the top 200 markers by coverage and / or those with a coverage greater than 1%; S5) Targeted methylation sequencing is performed on the biomarker set B' in the body fluid sample of the cancer type and the body fluid sample of one or more other cancer types to obtain the coverage E2 of each biomarker in the cancer type, and the biomarker set D is obtained based on the coverage E2. The process of obtaining the coverage E2 of each biomarker in the cancer type includes: selecting reads with a methylation rate of 0.5 or higher and covering more than 3 CpG sites as extremely hypermethylated reads; the proportion of an extremely hypermethylated read of a biomarker to the total reads of that biomarker as the extreme hypermethylated read ratio of that biomarker; and binarying the extreme hypermethylated read ratios of each biomarker in the obtained cancer type body fluid samples based on threshold quantiles to determine the coverage E2 of each biomarker. In step S5), the determination of the threshold quantile includes: statistically analyzing the extreme hypermethylation read ratios of each marker in body fluid samples of one or more other cancer types, and using the 90% to 99% quantile of the extreme hypermethylation read ratios as the threshold quantiles of each marker. In step S5), obtaining the marker set D based on coverage E2 includes: selecting the top 200 markers in terms of coverage and / or those with a coverage greater than 1%. S6) Merge the set of markers C and the set of markers D to remove duplicates to obtain the set of markers; The method further includes: using bodily fluid samples of the cancer type and the non-cancer bodily fluid samples, employing an iterative algorithm to screen for sensitive methylation biomarkers from biomarker set A', and adding them to biomarker set C. The step of using an iterative algorithm to screen for sensitivity methylation biomarkers from biomarker set A' includes: M1) The complementarity gain score Qi for each newly added marker i is calculated using the following formula: Where the weighting coefficient W1 is 1, the weighting coefficient W2 is -1, the weighting coefficient W3 is -2, Nc is the cancer body fluid sample in the cancer body fluid cohort of biomarker set A where the proportion of positive biomarkers is lower than the threshold P1, Nh1 is the non-cancer body fluid sample in the non-cancer body fluid sample cohort of biomarker set A where the proportion of positive biomarkers is higher than the threshold P2 but lower than the threshold P3, Nh2 is the non-cancer body fluid sample in the non-cancer body fluid sample cohort of biomarker set A where the proportion of positive biomarkers is higher than the threshold P3, Rin represents the extreme hypermethylation read ratio of the newly added biomarker i in sample n, and Ci represents the threshold quantile of the newly added biomarker i. The positive biomarker is a biomarker with an extremely high methylation read ratio higher than the threshold quantile; The threshold P1 is the 50%–70% percentile of the proportion of positive markers in cancer body fluid samples; The threshold P2 is the 85%–93% percentile of the proportion of positive markers in non-cancerous body fluid samples; The threshold P3 is the 93%–96% percentile of the proportion of positive markers in non-cancerous body fluid samples; M2) Based on the complementary gain score Qi, the marker with the highest score in each iteration is selected as the new sensitivity marker; and M3) Two conditions are set as iteration termination conditions in the iteration process: ① When Qi is less than or equal to 0, the iteration process of adding new biomarkers is stopped. ② The 90% to 99% percentile of the proportion of positive biomarkers in non-cancer samples is used as the threshold. When the number of cancer body fluid samples above the threshold after adding biomarkers is less than the number of cancer body fluid samples above the threshold before adding biomarkers, the iteration process is stopped.
2. The screening method according to claim 1, characterized in that, In step S3), a differential analysis is performed on the methylation regions of the cancer tissue sample of the stated cancer type and the tissue sample of another cancer type; and / or, a differential analysis is performed on the methylation regions of the cancer tissue sample of the stated cancer type one by one or simultaneously with the tissue samples of multiple other cancer types.
3. The screening method according to claim 1, characterized in that, In steps S2) and S3), the screening of methylation regions where the average methylation rate difference of the cancer type is greater than a threshold fraction includes: screening methylation regions where the average methylation rate difference of the cancer tissue sample is greater than 0.
1.
4. The screening method according to claim 1, characterized in that, In step S2), the filtering step includes: selecting a set A' of markers from the marker set A that has an average methylation rate of less than 0.5 in the region of the 85% to 99.9% percentile in non-cancer body fluid samples.
5. The screening method according to claim 1, characterized in that, In step S3), the filtering step includes: selecting from the biomarker set B a' that has an average methylation rate of less than 0.5 in the region of the 85% to 99.9% percentile in tissue samples of one or more other cancer types.
6. The screening method according to claim 1, characterized in that, The binary processing described in steps S4) and S5) includes: if the extreme hypermethylation read ratio of a marker in a cancer fluid sample is higher than the threshold quantile, then the marker is set to 1; otherwise, it is set to 0. The coverage is the proportion of cancerous fluid samples with a binary value of 1 for each marker to the total number of cancerous blood samples.
7. The screening method according to any one of claims 1 to 6, characterized in that, The cancer type is selected from any one of liver cancer, lung cancer, colorectal cancer, stomach cancer, esophageal cancer, pancreatic cancer, breast cancer, and ovarian cancer; and / or The non-cancerous samples include adjacent tissue or body fluid samples; and / or The fluid samples include: saliva or saliva supernatant, whole blood, serum, plasma, milk, urine, lumbar spine or ventricular CSF, lymph, prostatic fluid, semen, sputum, diluted fecal supernatant, tears, bronchoalveolar lavage fluid, sputum, pus, nasopharyngeal swabs, oral swabs, cerebrospinal fluid, pleural effusion, peritoneal fluid, amniotic fluid, peritoneal fluid, aqueous humor, vitreous humor, diluted vaginal discharge supernatant; and / or The tissue samples include tissue and paraffin sections.
8. An apparatus for screening methylation markers, used to implement the screening method for methylation markers according to any one of claims 1 to 7.
9. The apparatus according to claim 8, characterized in that, The device includes: a module for dividing methylation regions; The methylation sequencing results, based on cancer tissue samples and non-cancer samples corresponding to cancer types, divide the genome into multiple methylation regions; The first methylation region differential analysis module is used to perform differential analysis on the methylation regions of cancer tissue samples and non-cancer samples of the cancer type, screen out methylation regions whose average methylation rate difference of the cancer type is greater than a threshold fraction relative to non-cancer samples, and obtain the biomarker set A after merging and deduplication. The second methylation region differential analysis module is used to perform differential analysis on the methylation regions of the cancer tissue sample of the stated cancer type and tissue samples of one or more other cancer types, screen out the methylation regions whose average methylation rate difference is greater than a threshold fraction relative to tissue samples of one or more other cancer types, and merge and remove duplicates to obtain the biomarker set B. The first coverage analysis module is used to obtain the coverage E1 of each biomarker in the cancer type based on the targeted methylation sequencing results of biomarker set A in body fluid samples and non-cancer body fluid samples of the cancer type, and obtain biomarker set C based on coverage E1. The second coverage analysis module is used to obtain the coverage E2 of each biomarker in the cancer type based on the targeted methylation sequencing results of biomarker set B in body fluid samples of the stated cancer type and one or more other cancer types, and to obtain biomarker set D based on coverage E2; and The marker set acquisition module is used to merge marker set C and marker set D to obtain the marker set after deduplication.
10. A reagent for detecting a combination of methylation markers, wherein the combination of methylation markers is obtained by the screening method according to any one of claims 1 to 7, and wherein the combination of methylation markers includes cancer detection methylation markers and organ origin methylation markers; in, Based on the sequence of the human reference genome Hg19, the methylation markers for cancer detection are shown in the table below; Based on the sequence of the human reference genome Hg19, the organ-derived methylation markers are as shown in the table below: 。 11. The reagent for detecting combinations of methylation markers according to claim 10, characterized in that, The cancer types mentioned are any one or more of the following: liver cancer, lung cancer, colorectal cancer, stomach cancer, esophageal cancer, pancreatic cancer, breast cancer, and ovarian cancer.
12. The reagent for detecting combinations of methylation markers according to claim 10, characterized in that, The reagent for detecting the combination of methylation markers is a primer pair, which uses the nucleotide sequence containing the combination of methylation markers as the target sequence for specific amplification of the target sequence.
13. The reagent for detecting combinations of methylation markers according to claim 10, characterized in that, The reagent for detecting the combination of methylation markers is a probe, which is used to specifically capture the combination of methylation markers.
14. A storage medium comprising a computer program that, when executed by a processor, implements the screening method for methylation markers according to any one of claims 1 to 7.
15. Use of the apparatus for screening methylation markers according to claim 8 or 9, and the reagent for detecting combinations of methylation markers according to any one of claims 10 to 13, in the preparation of a product for cancer detection, wherein the cancer is one or more of liver cancer, lung cancer, colorectal cancer, gastric cancer, esophageal cancer, pancreatic cancer, breast cancer, and ovarian cancer.
16. The use according to claim 15, characterized in that, The cancer detection includes those used for early cancer screening, diagnosis, or prognosis prediction.
17. The use according to claim 16, characterized in that, The diagnosis includes auxiliary diagnoses.
Citation Information
Patent Citations
CtDNA methylation segment marker for diagnosing gastric cancer and predicting prognosis of gastric cancer
CN118086490A
Compositions and methods for identifying cell types
US20230212674A1