Methods and systems for evaluating tumor formation risk and tumor tissue origin.
A cost-effective and precise method for evaluating cancer risk and tumor tissue origin is achieved by segmenting DMRs and using machine learning to correct for confounding variables, addressing the limitations of existing DNA methylation sequencing.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- GUANGZHOU BURNING ROCK DX CO LTD
- Filing Date
- 2022-11-02
- Publication Date
- 2026-07-29
AI Technical Summary
Existing DNA methylation sequencing methods, such as WGBS, are costly and prone to DNA damage, and the discovery of cancer-related differentially methylated regions (DMRs) is challenging due to population heterogeneity and non-specific methylation changes, making it difficult to establish accurate tissue-of-origin models for cancer detection.
A method using DNA or RNA oligonucleotide sequences to capture methylated variant regions, segmenting DMRs based on sequencing coverage and methylation levels, and employing machine learning models to evaluate tumorigenesis risk and tumor tissue origin, incorporating a regularization term to correct for confounding variables.
This approach provides a low-cost, high-precision method for predicting cancer risk and determining tumor tissue origin with accuracy, achieving 98% tissue tracking accuracy in cross-validation and reducing the influence of age and other confounding factors.
Smart Images

Figure 0007897416000104 
Figure 0007897416000105 
Figure 0007897416000106
Abstract
Description
Technical Field
[0001] This application relates to the field of biomedicine, and specifically, to methods and systems for assessing the risk of tumor formation and the origin of tumor tissue.
Background Art
[0002] DNA methylation is known to play an important role in the regulation of gene expression. Abnormal DNA methylation markers have been reported in the occurrence and development of various diseases including cancer. As a high-resolution and high-throughput technology, DNA methylation sequencing is increasingly recognized for its role in cancer screening, diagnosis, and monitoring. Whole-genome bisulfite sequencing (WGBS, w hole g enome b isulfite s sequencing) is the gold standard for methylation sequencing, but it is difficult to apply clinically due to severe damage to DNA during processing and the too high sequencing cost. More importantly, most regions of the human genome are inactive during the occurrence and development of cancer, and mutations related to cancer tend to concentrate in specific regions such as CpG islands, thus providing an opportunity suitable for targeted sequencing.
[0003] However, the discovery and screening of cancer-related differentially methylated regions (DMRs) is challenging. This is because population heterogeneity, including disease and age status, can lead to non-specific changes in methylation profiles, and these non-cancerous but abnormal signals need to be processed during the DOC (Detection of Cancer) modeling process for cancer detection. Finally, establishing tissue-of-origin (TOO) models for applications in detecting multiple cancer types will be crucial in tracking possible organs of origin for cancer mutations, determining downstream diagnostic and treatment pathways, and saving healthcare costs. [Overview of the Initiative] [Problems that the invention aims to solve]
[0004] This application establishes a low-cost, highly accurate method that uses DNA or RNA oligonucleotide sequences to capture methylated variant regions in many different cancers and specific methylated signature regions in various organs, to determine the presence of tumor fractions (ctDNA) in free DNA (cfDNA) in the blood, and to evaluate the correlation between samples and tumor tissue origins. [Means for solving the problem]
[0005] In one embodiment, the present application provides a method for evaluating the correlation between a test sample and tumorigenesis risk and / or tumor tissue origin, comprising: (1) a step of segmenting methylation variable regions (DMRs): determining a plurality of target DMRs to be used for evaluation based on the depth of sequencing coverage of methylation sites and / or the degree of difference in methylation levels of adjacent methylation sites; (2) a step of evaluating tumorigenesis risk: evaluating the correlation between the test sample and tumorigenesis risk based on the methylation levels of the target DMRs of the test sample; and (3) an optional step of evaluating tumor tissue origin: evaluating the correlation between the test sample and tumor tissue origin based on the methylation levels of the target DMRs of the test sample.
[0006] In one embodiment, the present application provides a method for determining a methylation variable region (DMR), the method comprising: dividing the methylation variable region (DMR); determining the methylation variable region (DMR) based on the depth of sequencing coverage of the methylation sites and / or the degree of difference in methylation levels of adjacent methylation sites.
[0007] In one embodiment, the present application provides a method for evaluating the correlation between a test sample and tumorigenesis risk, comprising the steps of: evaluating the tumorigenesis risk; evaluating the correlation between a test sample and tumorigenesis risk based on the methylation level of the DMR of the test sample, comprising the steps of reducing the influence of the subject's age factor on the evaluation result, wherein the test sample is derived from the subject.
[0008] In one embodiment, the present application provides a method for evaluating the correlation between a test sample and tumor tissue origin, comprising the steps of: evaluating the correlation between a test sample and tumor tissue origin based on the methylation level of the DMR in the test sample, using a multi-class classification method and a logistic regression method.
[0009] In one embodiment, the present application provides a storage medium that stores a program capable of performing the method described herein.
[0010] In one embodiment, the present application provides an apparatus comprising a storage medium described herein and optionally a processor coupled to the storage medium, wherein the processor is configured to implement the method described herein based on the execution of a program stored in the storage medium.
[0011] In one embodiment, the present application provides a system for evaluating the correlation between a test sample and the risk of tumor formation and / or tumor tissue origin, comprising: (1) a methylation variable region segmentation module: a module for determining a plurality of target DMRs to be used for evaluation based on the depth of coverage of sequencing of methylation sites and / or the degree of difference in methylation levels of adjacent methylation sites; (2) a tumor formation risk evaluation module: a module for evaluating the correlation between a test sample and the risk of tumor formation based on the methylation level of the target DMRs of the test sample; and (3) optionally, a tumor tissue origin evaluation module: a module for evaluating the correlation between a test sample and tumor tissue origin based on the methylation level of the target DMRs of the test sample.
[0012] In one embodiment, the present application provides a system for determining a methylation variable region (DMR), comprising: a methylation variable region DMR partitioning module: a module for determining a methylation variable region DMR based on the depth of sequencing coverage of methylation sites and / or the degree of difference in methylation levels of adjacent methylation sites.
[0013] In one embodiment, the present application provides a system for evaluating the correlation between a test sample and tumorigenesis risk, comprising: a tumorigenesis risk evaluation module: a module for evaluating the correlation between a test sample and tumorigenesis risk based on the methylation level of the DMR of the test sample, and a module for reducing the influence of the subject's age factor on the evaluation results, wherein the test sample is derived from the subject.
[0014] one In this embodiment, the present application provides a system for evaluating the correlation between a test sample and tumor tissue origin, comprising: a tumor tissue origin evaluation module: a module for evaluating the correlation between a test sample and tumor tissue origin based on the methylation level of the DMR of the test sample using a multi-class classification method and a logistic regression method. [Effects of the Invention]
[0015] This application contributes to accurate prediction and assessment of the risk of various cancers by providing a low-cost, high-precision method.
[0016] Those skilled in the art will readily be able to infer other aspects and advantages of this application from the following detailed description. The following detailed description shows and describes only exemplary embodiments of this application. As will be apparent to those skilled in the art, the content of this application will enable those skilled in the art to make modifications to the specific embodiments disclosed without departing from the spirit and scope of the invention to which this application relates. Accordingly, the accompanying drawings and descriptions of this application are illustrative and not intended to limit. [Brief explanation of the drawing]
[0017] The specific features of the invention relating to this application are set forth in the attached claims. The features and advantages of the invention relating to this application can be better understood by referring to the exemplary embodiments and accompanying drawings described in detail below. A brief description of the accompanying drawings is given below.
[0018] [Figure 1] Shows an exemplary situation (a theoretical exemplary demonstration and not intended to represent actual sequencing situations). [Figure 2] Shows another exemplary situation (a theoretical exemplary demonstration and not intended to represent actual sequencing situations). [Figure 3A-3C] Shows another exemplary situation (a theoretical exemplary demonstration and not intended to represent actual sequencing situations). [Figure 4] Shows that in five-fold cross-validation, tissue tracking accuracy of 98% (95% CI: 96 - 99%) can be achieved. [Figure 5] Shows the control results of the weight allocation of the confounding correlation features in the Salmon-DOC model of the present application. [Figures 6A-6F] Shows that the Salmon-DOC model of the present application in the tumor group model can efficiently achieve the detection of different stages of six cancer types. [Figures 7A-7F] Shows that in the healthy group, the Salmon-DOC model of the present application overcomes the weakness of methylation false positives that increases with age and maintains balance in each age group (age on the horizontal axis, cancer probability score of the model on the vertical axis). [Figures 8A-8D] Shows that the tracking accuracy of the Salmon-TOO two-layer model of the present application is better than that of the one-layer model in both cross-validation and independent validation. [Figure 9] Shows the results of tissue tracking evaluation obtained based on 103 TOO-related DMR regions.
Mode for Carrying Out the Invention
[0019] Hereinafter, embodiments of the present invention will be described with specific specific embodiments, but other advantages and effects of the present invention can be easily understood by those skilled in the art from the content disclosed in this specification.
[0020] Definition of Terms In this application, the terms “second-generation gene sequencing (NGS),” “high-throughput sequencing,” or “next-generation sequencing” generally refer to second-generation high-throughput sequencing technologies and subsequent higher-throughput sequencing methods. Next-generation sequencing platforms include, but are not limited to, existing sequencing platforms such as Illumina. As sequencing technologies continue to evolve, it will be understood by those skilled in the art that other sequencing methods and apparatus may also be employed for use in the methods of the present invention. For example, second-generation gene sequencing may offer advantages such as high sensitivity, high throughput, high sequencing depth, or low cost. Depending on their development history, influence, different sequencing principles, and technologies, the following are some of the main types of sequencing methods: Examples include Massively Parallel Signature Sequencing (MPSS), Polony Sequencing, 454 pyrosequencing, Illumina (Solexa) sequencing, Ion Semiconductor Sequencing, DNA Nano-ball Sequencing, and Complete Genomics' DNA nanoarray and probe-anchoring ligation combined sequencing method. The aforementioned second-generation sequencing enables detailed and comprehensive analysis of the transcriptome and genome of a single species and is therefore also called deep sequencing. For example, the method of this application can also be applied to first-generation gene sequencing, second-generation gene sequencing, third-generation gene sequencing, or single-molecule sequencing (SMS).
[0021] In this application, the term "tested sample" generally refers to the sample being tested. For example, the presence or absence of modification status can be detected in one or more gene regions of the tested sample.
[0022] In this application, the terms “polynucleotide,” “nucleotide,” “nucleic acid,” and “oligonucleotide” are used interchangeably. These refer to polymeric forms of nucleotides (deoxyribonucleotides or ribonucleotides) of any length, or analogs thereof. Polynucleotides can have any three-dimensional structure and can perform any known or unknown function. The following are non-limiting examples of polynucleotides: coding or non-coding regions of genes or gene fragments, loci (gene loci) identified according to linkage analysis, exons, introns, messenger RNA (mRNA), transporter RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA having any sequence, isolated RNA having any sequence, nucleic acid probes, primers, and linkers. Polynucleotides may include one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs.
[0023] In this application, the term "methylation" generally refers to the methylation state of the gene fragment, nucleotide, or base of the gene of this application. For example, the DNA fragment containing the gene of this application may be methylated on one or more strands. For example, the DNA fragment containing the gene of this application may be methylated at one site or at multiple sites.
[0024] In this application, the term “human reference genome” generally refers to a human genome capable of performing a reference function in gene sequencing. Information regarding the human reference genome can be found on UCSC. The human reference genome is available in different versions, e.g., hg19, GRCH37, or ensembl 75.
[0025] In this application, the term “machine learning model” generally refers to a system or program instructions and / or a set of data arranged to implement an algorithm, process, or mathematical model. In this application, the algorithm, process, or mathematical model can be evaluated based on a given input and provide a desired output. In this application, the parameters of the machine learning model do not have to be explicitly programmed, and in the conventional sense, the machine learning model does not have to be explicitly designed to follow specific rules to provide a desired output for a given input. For example, the use of the machine learning model may mean that the machine learning model and / or a set of data structures / rules as a machine learning model is trained by a machine learning algorithm.
[0026] In this application, the term "includes" generally means including the features that are explicitly specified, but does not mean excluding other factors.
[0027] In this application, the term "about" generally means that the value varies within a range of ±0.5 to 10% of the specified value, for example, within a range of ±0.5%, ±1%, ±1.5%, ±2%, ±2.5%, ±3%, ±3.5%, ±4%, ±4.5%, ±5%, ±5.5%, ±6%, ±6.5%, ±7%, ±7.5%, ±8%, ±8.5%, ±9%, ±9.5%, or ±10% of the specified value.
[0028] To detect six types of cancer with high incidence and mortality rates, including lung cancer, colorectal cancer, liver cancer, ovarian cancer, pancreatic cancer, and esophageal cancer, this application employs a novel algorithm that combines a public database (TCGA) with in-house data mining to simultaneously compare genomic methylation variants and their spatial locations, screening a total of 2,536 differentially methylated regions (DMRs) that are highly correlated with cancer.
[0029] Summary of the Invention In one embodiment, the present application provides a method for evaluating the correlation between a test sample and tumorigenesis risk and / or tumor tissue origin, which may include: (1) a step of segmenting methylation variable regions (DMRs): determining a plurality of target DMRs to be used for evaluation based on the depth of sequencing coverage of methylation sites and / or the degree of difference in methylation levels of adjacent methylation sites; (2) a step of evaluating tumorigenesis risk: evaluating the correlation between the test sample and tumorigenesis risk based on the methylation levels of the target DMRs of the test sample; and (3) an optional step of evaluating tumor tissue origin: evaluating the correlation between the test sample and tumor tissue origin based on the methylation levels of the target DMRs of the test sample. For example, the method for evaluating the correlation between a test sample and tumorigenesis risk and / or tumor tissue origin according to this application may include: (1) determining a number of target DMRs to be used for evaluation based on the depth of sequencing coverage of methylation sites and / or the degree of difference in methylation levels of adjacent methylation sites; (2) evaluating the correlation between the test sample and tumorigenesis risk based on the methylation levels of the target DMRs in the test sample; and (3) optionally evaluating the correlation between the test sample and tumor tissue origin based on the methylation levels of the target DMRs in the test sample.
[0030] In one embodiment, the present application provides a method for determining a methylation variable region (DMR), which may include the steps of: dividing the methylation variable region (DMR); determining the methylation variable region (DMR) based on the depth of sequencing coverage of the methylation sites and / or the degree of difference in methylation levels of adjacent methylation sites.
[0031] In one embodiment, the present application provides a method for evaluating the correlation between a test sample and tumorigenesis risk, comprising the steps of: evaluating tumorigenesis risk; evaluating the correlation between a test sample and tumorigenesis risk based on the methylation level of the DMR of the test sample, comprising reducing the influence of the subject's age factor on the evaluation result, wherein the test sample is derived from the subject.
[0032] In one embodiment, the present application provides a method for evaluating the correlation between a test sample and tumor tissue origin, comprising the steps of: evaluating the correlation between the test sample and tumor tissue origin based on the methylation level of the DMR in the test sample by a multi-class classification method and a logistic regression method.
[0033] For example, the method may include determining the DMR based on the depth of sequencing coverage of the methylation site and the degree of difference in methylation levels between the methylation site and its adjacent methylation sites. For example, the degree of difference in methylation levels may refer to the difference in methylation levels. For example, the degree of difference in methylation levels may refer to the absolute value of the difference in methylation levels. For example, this application can determine DMR regions having substantially the same methylation level based on the degree of difference in methylation levels between the methylation site and its adjacent methylation sites. For example, this application can more accurately segment DMR regions by the depth of sequencing coverage of the methylation site. For example, data information from sites with deeper coverage is more reliable.
[0034] For example, the method may include determining the absolute value of the difference in methylation levels between a methylation site and its adjacent methylation sites, and determining, based on the absolute value of the difference, whether the methylation site and its adjacent methylation sites are divided into the same DMR. For example, the method may include determining the weight of the absolute value of the difference, the weight of the absolute value of the difference is determined based on the sequencing coverage depth of the methylation site. For example, JPEG0007897416000001.jpg16170
[0035] For example, the method may include determining the absolute value of the difference in methylation levels between a methylation site and an adjacent methylation site, and determining the weight of the absolute value of the difference in methylation levels, the weight of the absolute value of the difference being determined based on the sequencing coverage depth of the methylation site.
[0036] JPEG0007897416000002.jpg31170JPEG0007897416000003.jpg16170JPEG0007897416000004.jpg38170
[0037] In the case of JPEG0007897416000005.jpg9170, it is determined that the methylation site and its adjacent methylation sites are divided into the same DMR.
[0038] For example, the method may further include determining the degree of variation in the methylation level of the DMR based on the difference in the degree of difference in methylation levels between the methylation sites of the DMR and the methylation sites at intermediate positions of the DMR. For example, the intermediate position means an intermediate position in terms of physical location. For example, if M is odd and the DMR has M methylation sites, the intermediate position may refer to the (M+1) / 2th methylation site from upstream to downstream. For example, if M is even and the DMR has M methylation sites, the intermediate position may refer to the M / 2nd or M / 2+1th methylation site from upstream to downstream.
[0039] For example, by determining the degree of variation between the degree of methylation differences at each methylation site in the candidate DMR and the degree of variation between the degree of methylation differences at intermediate methylation sites, a more preferable DMR can be screened among the candidate DMRs.
[0040] Determining a DMR of less than approximately 1 is used to assess the correlation between the test sample and the risk of tumorigenesis and / or tumor tissue origin.
[0041] For example, the method may include a step of evaluating whether the test sample has a risk of tumorigenesis using a binary classification model based on the methylation level of the DMR of the test sample, the evaluation method reduces the influence of the subject's age factor on the evaluation results of the correlation between the test sample and tumorigenesis risk and / or tumor tissue origin, and the test sample originates from the subject.
[0042] For example, the binary classification model may include a support vector machine (SVM) model. For example, the method may include introducing a penalty term based on the age factor in the SVM model. For example, the method may include introducing a penalty term based on the age factor in the SVM model using the Hilbert-Schmidt independence criterion. For example, all of the methods for introducing a penalty term in machine learning in this application can be used in this application to reduce the influence of the age factor.
[0043] For example, the above method may include performing machine learning training on training samples in which the presence or absence of tumor formation is known, according to the following formula. JPEG0007897416000008.jpg29170
[0044] The following formula is used to determine the training parameters: JPEG0007897416000009.jpg42170JPEG0007897416000010.jpg33170JPEG0007897416000011.jpg40170JPEG0007897416000012.jpg16170
[0045] For example, the method may include the steps of determining classification probabilities by a multi-class classification method based on the methylation level of the DMR of the test sample, and fitting the classification probabilities by logistic regression to evaluate the correlation between the test sample and tumor tissue origin. For example, the method may determine classification probabilities by pairwise voting. For example, the method may determine classification probabilities by various multi-class classification methods in the field. For example, the method may fit the classification probabilities by multiple linear regression (MLR).
[0046] For example, the method may include performing a regression analysis on training samples of known tissue origin according to the following formula: JPEG0007897416000013.jpg41170JPEG0007897416000014.jpg35170JPEG0007897416000015.jpg43170
[0047] For example, the method corrects for the tissue origin of the training sample based on the probability of tumor formation in the sample. For example, the method may include the step of performing the correction before obtaining classification probabilities by pairwise voting. For example, the method may include performing the correction after obtaining classification probabilities by pairwise voting and before performing the multiple linear regression analysis. For example, the method may include performing the correction based on a pseudo-likelihood estimation method.
[0048] For example, the method may include performing the correction according to the following formula, JPEG0007897416000016.jpg40170
[0049] For example, by maximizing the expected value of an equation and determining the weights, we can correct for tissue-derived classes based on whether or not a tumor has formed in the sample. For instance, evaluating samples with tumor formation makes the information derived from that tissue more reliable.
[0050] In one embodiment, the present application provides a system for evaluating the correlation between a test sample and the risk of tumorigenesis and / or tumor tissue origin, which may include: (1) a methylation variable region segmentation module: a module for determining a plurality of target DMRs to be used for evaluation based on the depth of coverage of sequencing of methylation sites and / or the degree of difference in methylation levels of adjacent methylation sites; (2) a tumorigenesis risk evaluation module: a module for evaluating the correlation between a test sample and the risk of tumorigenesis based on the methylation levels of the target DMRs of the test sample; and (3) optionally, a tumor tissue origin evaluation module: a module for evaluating the correlation between a test sample and tumor tissue origin based on the methylation levels of the target DMRs of the test sample.
[0051] In one embodiment, the present application provides a system for determining methylation variable region DMR, the system comprising a methylation variable region DMR partitioning module: a module for determining methylation variable region DMR based on the depth of sequencing coverage of methylation sites and / or the degree of difference in methylation levels of adjacent methylation sites.
[0052] In one embodiment, the present application provides a system for evaluating the correlation between a test sample and tumorigenesis risk, comprising: a tumorigenesis risk evaluation module: a module for evaluating the correlation between a test sample and tumorigenesis risk based on the methylation level of the DMR of the test sample, and a module for reducing the influence of the subject's age factor on the evaluation results, wherein the sample is derived from the subject.
[0053] In one embodiment, the present application provides a system for evaluating the correlation between a test sample and tumor tissue origin, comprising: a tumor tissue origin evaluation module: a module that evaluates the correlation between a test sample and tumor tissue origin based on the methylation level of the DMR of the test sample using a multi-class classification method and a logistic regression method.
[0054] For example, the system may include determining the DMR based on the depth of sequencing coverage of the methylation sites and the degree of difference in methylation levels between the methylation site and its adjacent methylation sites. For example, the degree of difference in methylation levels may refer to the difference in methylation levels. For example, the degree of difference in methylation levels may refer to the absolute value of the difference in methylation levels. For example, this application can determine DMR regions having substantially the same methylation level based on the degree of difference in methylation levels between the methylation site and its adjacent methylation sites. For example, this application can more accurately segment DMR regions by the depth of sequencing coverage of the methylation sites. For example, data information from sites with deeper coverage is more reliable.
[0055] For example, the system may include determining the absolute value of the difference in methylation levels between a methylation site and an adjacent methylation site, and determining, based on the absolute value of the difference, whether the methylation site and its adjacent methylation sites are divided into the same DMR. For example, the system may include determining the weight of the absolute value of the difference, the weight of the absolute value of the difference being determined based on the sequencing coverage depth of the methylation site. JPEG0007897416000017.jpg18170
[0056] For example, the system may include determining the absolute value of the difference in methylation levels between a methylation site and an adjacent methylation site, and determining the weight of the absolute value of the difference in methylation levels, the weight of the absolute value of the difference being determined based on the depth of sequencing coverage of the methylation site.
[0057] JPEG0007897416000018.jpg30170JPEG0007897416000019.jpg15170JPEG0007897416000020.jpg45170
[0058] In the case of JPEG0007897416000021.jpg7170, it is determined that the methylation site and its adjacent methylation sites are divided into the same DMR.
[0059] For example, the system further includes determining the degree of variation in the methylation level of the DMR based on the difference in the degree of difference in methylation levels between the methylation sites of the DMR and the methylation sites at intermediate positions of the DMR. For example, the intermediate position means an intermediate position in terms of physical location. For example, if M is odd and the DMR has M methylation sites, the intermediate position may refer to the (M+1) / 2th methylation site from upstream to downstream. For example, if M is even and the DMR has M methylation sites, the intermediate position may refer to the M / 2nd or M / 2+1th methylation site from upstream to downstream.
[0060] For example, by determining the degree of variation between the degree of methylation differences at each methylation site in the candidate DMR and the degree of variation between the degree of methylation differences at intermediate methylation sites, a more preferable DMR can be screened among the candidate DMRs.
[0061] Determining a DMR of less than approximately 1 is used to assess the correlation between the test sample and the risk of tumorigenesis and / or tumor tissue origin.
[0062] For example, the system may include evaluating whether a test sample has a risk of tumorigenesis using a binary classification model based on the methylation level of the DMR in the test sample, the evaluation system reduces the influence of the subject's age factor on the evaluation results of the correlation between the test sample and tumorigenesis risk and / or tumor tissue origin, and the test sample originates from the subject.
[0063] For example, the binary classification model may include a support vector machine (SVM) model. For example, the system may include introducing a penalty term based on the age factor in the SVM model. For example, the system may include introducing a penalty term based on the age factor in the SVM model using the Hilbert-Schmidt independence criterion. For example, all the methods for introducing a penalty term in machine learning in this application can be used in this application to reduce the influence of the age factor.
[0064] For example, the system may include performing machine learning training on training samples in which the presence or absence of tumor formation is known, according to the following formula. JPEG0007897416000024.jpg29170
[0065] The following formula is used to determine the training parameters: JPEG0007897416000025.jpg48170JPEG0007897416000026.jpg35170JPEG0007897416000027.jpg45170JPEG0007897416000028.jpg16170
[0066] For example, the system may include determining classification probabilities by a multi-class classification method based on the methylation level of the DMR of the test sample, and fitting the classification probabilities by logistic regression to evaluate the correlation between the test sample and tumor tissue origin. For example, the system may determine classification probabilities by pairwise voting. For example, the system may determine classification probabilities by various multi-class classification method modules in the art. For example, the system may fit the classification probabilities by multiple linear regression (MLR).
[0067] For example, the system may include performing a regression analysis on training samples of known tissue origin according to the following formula: JPEG0007897416000029.jpg43170JPEG0007897416000030.jpg26170JPEG0007897416000031.jpg34170JPEG0007897416000032.jpg16170
[0068] For example, the system corrects for the tissue origin of the training sample based on the probability of tumor formation in the sample. For example, the system may include performing the correction before obtaining classification probabilities by pairwise voting. For example, the system may include performing the correction after obtaining classification probabilities by pairwise voting and before performing the multiple linear regression analysis. For example, the system may include performing the correction based on a pseudo-likelihood estimation method.
[0069] For example, the system may include performing corrections according to the following formula: JPEG0007897416000033.jpg41170
[0070] For example, by maximizing the expected value of an equation and determining the weights, we can correct for tissue-derived classes based on whether or not a tumor has formed in the sample. For instance, evaluating samples with tumor formation makes the information derived from that tissue more reliable.
[0071] In one embodiment, the present application provides a storage medium recording a program capable of performing the method described herein. For example, the non-volatile computer-readable storage medium may include floppy disks, flex disks, hard disks, solid-state storage (SSS) (e.g., solid-state drives (SSDs)), solid-state cards (SSCs), solid-state modules (SSMs)), enterprise flash drives, magnetic tapes, or any other non-transient magnetic media. The non-volatile computer-readable storage medium may also include punch cards, paper tapes, photomarker sheets (or any other physical medium having a hole pattern or other optically identifiable markings), compact disc read-only memory (CD-ROM), rewritable compact discs (CD-RWs), digital multipurpose discs (DVDs), Blu-ray discs (BDs), and / or any other non-transient optical media.
[0072] In one embodiment, the present application provides an apparatus comprising a storage medium described herein and optionally a processor coupled to the storage medium, wherein the processor is configured to implement the method described herein based on the execution of a program stored in the storage medium.
[0073] <Example 1> Exemplary bisulfite-treated second-generation sequencing was performed on the samples, and the resulting sequencing data included methylation levels of methylation sites (CpGs) and sequencing coverage depth. Optionally, denoising was performed on genomic methylation signals (CpGs) and noise regions (CHH / CHGs). Next, p-values obtained by weighted logistic regression were calculated for the "tumor" (C) and "normal" (N) groups. The explanatory variable for the logistic regression was a continuous variable, i.e., the methylation level of each CpG site, and the response variable was a binary output, i.e., (0,1) corresponding to C and N. The weighted logistic regression tested the distinction between C and N at each CpG site, with the null hypothesis being that the difference between C and N at that CpG site is not statistically significant. The weights were determined based on the coverage depth of each CpG site.
[0074] DMR splitting The method for partitioning each region of the DMR is determined based on the methylation level of the methylation site CpG and the sequencing coverage depth. Specifically, the methylation level of the methylation site CpG and the sequencing coverage depth are calculated according to the following formula. JPEG0007897416000034.jpg39170JPEG0007897416000035.jpg30170JPEG0007897416000036.jpg23170 The larger the value, the higher the similarity of methylation levels between adjacent CpG sites within the same group.
[0075] Figure 1 shows an illustrative scenario (a theoretical, illustrative demonstration, and not intended to represent actual sequencing conditions).
[0076] For the first CpG site in this region, samples A and B each obtained coverage of 500 valid sequences, while sample C obtained coverage of 200 valid sequences. In sample A, the methylation level of this CpG site is 0.2. The methylation level of the second CpG site in sample A is 0. For the three samples, the coverage depth parameter P for the first CpG site in this group was calculated to be 0.617. JPEG0007897416000037.jpg33170
[0077] Figure 2 shows another illustrative scenario (a theoretical illustrative demonstration and not intended to represent actual sequencing conditions).
[0078] If the above samples are replaced with A, B, and D (where sample D obtained 400 effective sequence coverage at the first CpG site), similarly, the methylation level of this CpG site in sample A is 0.2. The methylation level of the second CpG site in sample A is 0. However, in this example, because the sequencing coverage depth of sample D increased, the coverage depth of the first CpG site in the group increased in the three samples. JPEG0007897416000038.jpg21170
[0079] Therefore, by introducing coverage depth for the CpG region using the method of the present invention, the accuracy of DMR region segmentation can be significantly improved.
[0080] JPEG0007897416000039.jpg54170
[0081] Figures 3A-3C show another exemplary scenario (a theoretical illustrative demonstration and not intended to represent actual sequencing conditions). Ten DMR regions JPEG0007897416000040.jpg11170
[0082] Of these, the calculation steps for the values within the DMR region shown in Group A are as follows: JPEG0007897416000041.jpg74170
[0083] JPEG0007897416000042.jpg27170
[0084] The DMR regions screened using this method include not only cancer mutation information for various cancer types but also tissue-specific features, and exhibit high segmentation effectiveness at region boundaries.
[0085] Figure 4 shows lung cancer ( L ung C Arcinoma (LC), colon cancer ( C olorectal C Arcinoma (CRC), liver cancer ( L i ver H Epatocellular Carcinoma (LIHC), Ovarian Cancer ( O Varian C Arcinoma (OVCA), pancreatic cancer ( P ancreatic A denocarcinoma (PAAD), esophageal cancer ( E sophageal C This study demonstrates that a tissue tracking accuracy of 98% (95% CI: 96% to 99%) can be achieved in quintuple cross-validation for six types of arcinoma (ESCA) cancers.
[0086] <Example 2> Cancer Assessment (DOC) Modeling The content of ctDNA in the blood varies greatly depending on the different stages of development of different cancers and is susceptible to experimental batch effects. In addition, methylation variants are associated with age, disease, race, etc., and if left unaddressed, these can affect the accuracy of classification models as confounding variables. This application employs a modeling method called Salmon, which first quantifies the bias caused by confounding variables (quantification can be performed by the Hilbert-Schmidt independence criterion, but is not limited to this), and then incorporates a regularization term into the model to correct for it, thereby improving the accuracy and generalization of the model.
[0087] Establishment of an algorithm JPEG0007897416000043.jpg25170JPEG0007897416000044.jpg20170
[0088] JPEG0007897416000045.jpg40170JPEG0007897416000046.jpg26170JPEG0007897416000047.jpg32170
[0089] We primarily use a support vector machine (SVM) as our classifier. The classification interface for JPEG0007897416000048.jpg28170 is determined by solving the following target equation: For inseparable data, a soft-margin support vector machine (soft-margin SVM) introduces a penalty term for training error, JPEG0007897416000050.jpg45170JPEG0007897416000051.jpg15170
[0090] JPEG0007897416000052.jpg21170JPEG0007897416000053.jpg38170JPEG0007897416000054.jpg14170
[0091] Figure 5 shows the results of controlling the weighting of confounding correlation features in the Salmon-DOC model of this application.
[0092] Each data point represents a blood sample used to construct the Salmon-DOC model. The horizontal axis represents the Confusing Factor of the corresponding sample, and the vertical axis represents the original uncorrected Variable Coef (Figure A) and the corrected Variable Coef (Figure B), respectively. Comparing the uncorrected and corrected values, it can be seen that the weights of confounding correlation features are controlled in Salmon-DOC.
[0093] Retrospective cohort data In this application, the accuracy of Salmon's binary classifier (cancer vs. non-cancer) was evaluated using retrospective clinical samples of six cancer types, divided into a training set and a validation set.
[0094] Figures 6A to 6F demonstrate that the Salmon-DOC model of this application can efficiently detect different stages of six cancer types in a tumor group model.
[0095] Figures 7A-7F show that in the healthy control group, the Salmon-DOC model of this application overcomes the weakness of conventional models, which suffer from methylation-induced false positives that increase with age, and maintains balance across all age groups (age on the horizontal axis, model cancer probability score on the vertical axis).
[0096] <Example 3> Organizational Tracking (TOO) Modeling Building the first layer of the TOO model The TOO model is essentially a multi-class classification problem, and calculating the probability of each class involves voting on pairwise results and selecting the result with the most votes. However, for the possible clinical applications of tissue tracking models, simply producing classification results is insufficient; it is necessary to generate classification probabilities in order to enable model assembly.
[0097] Therefore, the first step in the Salmon-TOO model in this application is to quantify the binary voting results. This quantification is performed by probability calculation. JPEG0007897416000055.jpg28170JPEG0007897416000056.jpg35170JPEG0007897416000057.jpg96170
[0098] Building the second layer of the TOO model The second layer of the Salmon-TOO model is to apply MLR to different classes.
[0099] JPEG0007897416000058.jpg42170JPEG0007897416000059.jpg83170
[0100] JPEG0007897416000060.jpg24170
[0101] JPEG0007897416000061.jpg22170JPEG0007897416000062.jpg23170
[0102] In the Salmon-DOC model, it is found that some cancer types are judged negative and others positive. Therefore, when performing follow-up modeling on this judgment, weight correction based on pseudo-likelihood estimation is applied to the tissue class. Taking binary logistic regression as an example, it can be interpreted as follows. JPEG0007897416000063.jpg23170
[0103] Retrospective cohort data All data from the retrospective cohort were randomly split into a training set and a validation set in a 1:1 ratio. First, follow-up evaluation results were obtained through cross-validation using the training set, during which model parameters were continuously optimized and finally locked. Finally, all data from the validation set were evaluated with the locked model to obtain follow-up results. In the training set of the follow-up model, the sample size for the six cancer types totaled 300 cases, with a relatively balanced number of cases for each cancer type and stage. There were 36 cases of lung cancer (4 / 12 / 5 / 15 cases in stages I-IV, respectively), 62 cases of colorectal cancer (8 / 18 / 18 / 18 / 18 cases in stages I-IV, respectively), 74 cases of liver cancer (25 / 14 / 22 / 13 cases in stages I-IV, respectively), 48 cases of ovarian cancer (1 / 4 / 38 / 5 cases in stages I-IV, respectively), 40 cases of pancreatic cancer (3 / 6 / 13 / 18 cases in stages I-IV, respectively), and 42 cases of esophageal cancer (5 / 10 / 15 / 12 cases in stages I-IV, respectively). The follow-up validation set consisted of a total of 224 samples, including 31 cases of lung cancer (stages I-IV with 4 / 5 / 12 / 10 cases each), 52 cases of colorectal cancer (stages I-IV with 7 / 15 / 13 / 17 cases each), 55 cases of liver cancer (stages I-IV with 17 / 11 / 20 / 7 cases each), 27 cases of ovarian cancer (stages I-IV with 3 / 4 / 8 / 12 cases each), 25 cases of pancreatic cancer (stages I-IV with 4 / 6 / 6 / 9 cases each), and 34 cases of esophageal cancer (stages I-IV with 4 / 7 / 8 / 15 cases each).
[0104] Figures 8A–8D demonstrate that the tracking accuracy of the Salmon-TOO 2-layer model in this application is superior to that of the 1-layer model in both cross-validation and independent validation.
[0105] Figures 8A and 8B show the follow-up evaluation results of cross-validation of the six cancer types in the training set of six cancers. Figure 8A shows the output after constructing only the first layer TOO model, with a follow-up accuracy of 0.87 (260 / 300). Incorporating the suboptimal follow-up results improves the accuracy to 0.93 (279 / 300). Figure 8B shows the output after adding the second layer MLR model based on the first layer TOO model, improving the follow-up accuracy to 0.90 (270 / 300). Incorporating the suboptimal follow-up results further improves the accuracy to 0.95 (284 / 300). Similarly, Figures 8C and 8D show the follow-up evaluation results of independent validation of the six cancer types in the validation set described above. Of these, Figure 8C shows the output after constructing only the first layer TOO model, with a tracking accuracy of 0.77 (173 / 224). When the suboptimal tracking results are incorporated, the accuracy becomes 0.87 (194 / 224). Figure 8D shows the output after adding the second layer MLR model based on the first layer TOO model, with the tracking accuracy increased to 0.84 (187 / 224). When the suboptimal tracking results are incorporated, the accuracy can be further increased to 0.89 (199 / 224).
[0106] In summary, these results show that the evaluation accuracy of the Salmon-TOO 2-layer tracking model in this application is superior to that of the 1-layer model in both cross-validation and independent validation of the training set.
[0107] <Example 4> DOC Cancer Detection Model Table 1A shows the 94 DMR regions used in the DOC cancer detection model. JPEG0007897416000064.jpg229170JPEG0007897416000065.jpg70170
[0108] Based on 94 DOC-correlated DMR regions, evaluation was performed on 100 healthy control samples and 318 cancer-positive samples from 6 different cancers in independent validation set 1. The overall sensitivity was 80.5% (256 / 318), and the overall specificity was 95% (95 / 100). The sensitivity for specific cancer types and stages, while maintaining specificity at the 90% level, is shown in the table below. JPEG0007897416000066.jpg186170
[0109] Next, repeated tests were performed, with 50 randomly selected from 94 DOC regions for each test. The sensitivity results for six cancer-positive samples over five repeated tests, while maintaining a specificity level of 90% (90 / 100), are shown in the table below. JPEG0007897416000067.jpg76170
[0110] <Example 5> TOO Organization Tracking Model Table 1B shows the 103 DMR regions used in the TOO tissue tracking model. JPEG0007897416000068.jpg124170JPEG0007897416000069.jpg208170
[0111] Based on 103 TOO-correlated DMR regions, follow-up evaluation of 473 cancer-positive samples from 6 types in independent validation set 2 showed an initial follow-up accuracy of 63.0% (298 / 473 cases). By incorporating suboptimal follow-up results, this accuracy can be increased to 71.5% (338 / 473 cases).
[0112] Figure 9 shows the results of the tissue follow-up assessment obtained based on 103 TOO-related DMR regions.
[0113] Next, four iterative tests were conducted, with 50 randomly selected from 103 TOO areas each time. The results of the tracking accuracy over the four evaluations are shown in the table below. JPEG0007897416000070.jpg46170
[0114] <Example 6> Simultaneous evaluation results of DOC and TOO for 222 DMRs: Table 1C shows the 222 DMR regions used in the DOC and TOO evaluation models. JPEG0007897416000071.jpg111170JPEG0007897416000072.jpg232170JPEG0007897416000073.jpg234170JPEG0007897416000074.jpg94170
[0115] In an independent validation set, with 222 markers, sensitivity and follow-up accuracy were calculated for 473 negative samples and 473 positive samples of six types of cancer, with a unified specificity of 95.1% (450 / 473). The evaluation results for tumor examination and tissue follow-up are shown in the table below. JPEG0007897416000075.jpg61170JPEG0007897416000076.jpg47170
[0116] The detailed description above is provided for illustrative purposes and as an example, and is not intended to limit the scope of the appended claims. Multiple variations of the embodiments currently enumerated in this application will be apparent to those skilled in the art and remain within the scope of the appended claims and their equivalent embodiments.
Claims
1. A method for evaluating the correlation between a tested in vitro sample and the risk of tumor formation and / or tumor tissue origin, (1) Steps for dividing the methylation variable region DMR: Determining the absolute value of the difference in methylation levels between a methylation site and an adjacent methylation site; determining the weight of the absolute value of the difference according to the sequencing coverage depth of the methylation site; and determining a plurality of target DMRs to be used for evaluation based on the absolute value of the difference and the weight of the absolute value of the difference. (2) Steps to assess tumorigenesis risk: A step to assess the correlation between a test sample and tumorigenesis risk via a binary classification model that defines cancer or non-cancer based on the methylation level of the target DMR in the test sample, wherein the binary classification model is a support vector machine (SVM) model that incorporates a penalty term based on the subject's age factor, the penalty term is used to reduce the influence of the subject's age factor on the assessment results, and the test sample is derived from the subject, (3) An optional step to evaluate tumor tissue origin: A step to evaluate the correlation between the test sample and tumor tissue origin by determining the classification probability using a multi-class classification method based on the methylation level of the target DMR of the test sample, and then fitting the classification probability using logistic regression. A method characterized by including the following.
2. It is determined that the methylation site and its adjacent methylation sites are divided into the same DMR. The method according to claim 1.
3. The method further includes determining the degree of variation in the methylation level of the DMR based on the difference in the degree of difference in methylation levels between the methylation site of the DMR and the methylation site at an intermediate position of the DMR, This represents the degree of difference in methylation levels at methylation sites in the intermediate position of the MR region.
4. The SVM model includes introducing a penalty term based on the age factor using the Hilbert-Schmidt independence criterion, The method according to claim 1.
5. This includes performing machine learning training on training samples that are known to have tumorigenesis or are known not to have tumorigenesis, according to the following formula: The following formula is used to determine the training parameters: The method according to claim 1.
6. The method determines the classification probability by pairwise voting, Preferably, the method involves fitting the classification probabilities by multiple linear regression (MLR), Preferably, the method further includes performing a regression analysis on training samples of known tissue origin according to the following formula: Preferably, the method corrects for the tissue origin of the training sample based on the probability of tumor formation in the sample, Preferably, the method includes obtaining classification probabilities by pairwise voting and then performing the correction before the multiple linear regression analysis. Preferably, the method includes performing the correction based on a pseudo-likelihood estimation method. Preferably, the method includes making corrections according to the following formula: This represents the probability of a ulcer forming. The method according to claim 1.
7. A storage medium recording a program capable of performing the method described in any one of claims 1 to 6, wherein the program is executed by a processor.
8. A system for evaluating the correlation between a tested in vitro sample and the risk of tumor formation and / or tumor tissue origin, (1) Methylation Variable Region Segmentation Module: A module for determining the absolute value of the difference in methylation levels between a methylation site and an adjacent methylation site, determining the weight of the absolute value of the difference according to the sequencing coverage depth of the methylation site, and determining a plurality of target DMRs to be used for evaluation based on the absolute value of the difference and the weight of the absolute value of the difference. (2) Tumor formation risk assessment module: A module for evaluating the correlation between a test sample and tumor formation risk using a binary classification model that defines cancer or non-cancer based on the methylation level of the target DMR in the test sample, wherein the binary classification model is a support vector machine (SVM) model that incorporates a penalty term based on the subject's age factor, the penalty term is used to reduce the influence of the subject's age factor on the evaluation results, and the test sample is derived from the subject, (3) Module for evaluating the origin of an arbitrary tumor tissue: A module for evaluating the correlation between the test sample and tumor tissue origin by determining the classification probability using a multi-level classification method based on the methylation level of the target DMR of the test sample, and fitting the classification probability using logistic regression. A system comprising the above, wherein each module is realized by a device comprising a processor and a storage medium coupled thereto.
9. The system according to claim 8.
10. The method further includes determining the degree of variation in the methylation level of the DMR based on the difference in the degree of difference in methylation levels between the methylation site of the DMR and the methylation site at an intermediate position of the DMR, The system according to claim 8.
11. The SVM model includes introducing a penalty term based on the age factor using the Hilbert-Schmidt independence criterion, The system according to claim 8.
12. This includes performing machine learning training on training samples that are known to have tumorigenesis or are known not to have tumorigenesis, according to the following formula: The system according to claim 8.
13. The system determines the classification probability by pairwise voting, Preferably, the system applies the classification probabilities by multiple linear regression (MLR), Preferably, the system further includes performing a regression analysis on training samples of known tissue origin according to the following formula: Preferably, the system corrects the tissue origin of the training sample based on the probability of tumor formation in the sample. Preferably, the system includes performing the correction after obtaining the classification probability by the pairwise voting and before performing the multiple linear regression analysis. Preferably, the system includes performing the correction based on a pseudo-likelihood estimation method. Preferably, the system includes making corrections according to the following formula: The system according to claim 8.