Methods and systems for assessing tumor formation risk and tumor tissue origin
By employing DNA/RNA oligonucleotide sequences and machine learning, the method addresses the challenges of high costs and non-specific changes in DNA methylation sequencing, enabling precise cancer risk assessment and tumor tissue origin determination.
Patent Information
- Application Number
- JP2025504286
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-01
- Filing Date
- 2022-11-02
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2042-11-02
AI Technical Summary
Current DNA methylation sequencing methods, such as whole-genome bisulfite sequencing, face challenges due to DNA damage and high costs, and struggle with non-specific methylation changes in heterogeneous populations, making it difficult to accurately detect cancer-related methylation variants and determine the tissue of origin for cancer mutations.
A method using DNA or RNA oligonucleotide sequences to identify methylation variable regions (DMRs) based on sequencing coverage and methylation level differences, combined with machine learning models to evaluate tumor formation risk and tissue origin, reducing the influence of age and other confounding factors.
Provides a low-cost, high-precision method for predicting cancer risk and determining the origin of tumor tissue, achieving accurate classification and reducing the impact of population heterogeneity and age-related variations.
Smart Images

Figure 2025524966000001_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of biomedicine, and specifically, to methods and systems for evaluating the risk of tumor formation and the origin of tumor tissue.
Background Art
[0002] DNA methylation is known to play an important role in the regulation of gene expression. Abnormal DNA methylation markers have been reported in the occurrence and development of various diseases including cancer. As a high-resolution and high-throughput technology, DNA methylation sequencing is playing an increasingly recognized role in cancer screening, diagnosis, and monitoring. Whole-genome bisulfite sequencing (WGBS, w hole g enome b isulfite s sequencing) is the gold standard for methylation sequencing, but due to severe damage to DNA during processing and too high sequencing costs, its clinical application has become difficult. More importantly, most regions of the human genome are inactive during the occurrence and development of cancer, and cancer-related mutations tend to concentrate in specific regions such as CpG islands, thus providing an opportunity suitable for targeted sequencing.
[0003] However, the discovery and screening of methylation variable regions (Differentially Methylated Regions: DMRs) related to cancer are challenging. This is because the heterogeneity of the population, including disease, age, and other statuses, can lead to non-specific changes in the methylation profile, and these non-cancerous but abnormal signals need to be processed during the DOC (Detection Of Cancer) modeling process for cancer detection. Finally, establishing a tissue of origin (TOO) model for the detection of multiple cancer types is an important aid in tracking the possible origin organs of cancer mutations, determining downstream diagnostic and treatment pathways, and saving medical costs.
Summary of the Invention
Problems to be Solved by the Invention
[0004] This application uses DNA or RNA oligonucleotide sequences to capture methylation variant regions in many different cancers and specific methylation signature regions in various organs, determine the presence of tumor fraction (ctDNA) in cell-free DNA (cfDNA) in the blood, and evaluate the correlation between the sample and the tumor tissue origin, establishing a low-cost and high-precision method.
Means for Solving the Problems
[0005] In one aspect, the present application is a method for evaluating the correlation between a test sample and the risk of tumor formation and / or origin from tumor tissue, comprising: (1) a step of dividing methylation variable regions (DMRs): determining a plurality of target DMRs to be used in the evaluation based on the depth of sequencing coverage of methylation sites and / or the degree of difference in methylation levels of adjacent methylation sites; (2) a step of evaluating the risk of tumor formation: evaluating the correlation between the test sample and the risk of tumor formation based on the methylation level of the target DMRs in the test sample; (3) optionally, a step of evaluating origin from tumor tissue: evaluating the correlation between the test sample and origin from tumor tissue based on the methylation level of the target DMRs in the test sample.
[0006] In one aspect, the present application provides a method for determining methylation variable regions (DMRs), the method comprising a step of dividing methylation variable regions (DMRs): determining methylation variable regions (DMRs) based on the depth of sequencing coverage of methylation sites and / or the degree of difference in methylation levels of adjacent methylation sites.
[0007] In one aspect, the present application is a method for evaluating the correlation between a test sample and the risk of tumor formation, the method comprising a step of evaluating the risk of tumor formation: evaluating the correlation between the test sample and the risk of tumor formation based on the methylation level of the DMRs in the test sample, the method further comprising a step of reducing the influence of the age factor of the subject on the evaluation result, wherein the test sample is derived from the subject.
[0008] In one aspect, the present application is a method for evaluating the correlation between a test sample and origin from tumor tissue, the method comprising a step of evaluating origin from tumor tissue: evaluating the correlation between the test sample and origin from tumor tissue based on the methylation level of the DMRs in the test sample by a multi-class classification method and a logistic regression method.
[0009] In one aspect, the present application provides a storage medium storing a program capable of executing the method described in the present application.
[0010] In one aspect, the present application provides an apparatus including the storage medium described in the present application and, optionally, a processor coupled to the storage medium, the processor being arranged to implement the method described in the present application based on execution of a program stored in the storage medium.
[0011] In one aspect, the present application provides a system for evaluating the correlation between a test sample and the risk of tumor formation and / or origin from tumor tissue, the system comprising: (1) a methylation variable region division module: a module for determining a plurality of target DMRs to be used in the evaluation based on the depth of sequencing coverage of methylation sites and / or the degree of difference in methylation levels of adjacent methylation sites; (2) a tumor formation risk evaluation module: a module for evaluating the correlation between the test sample and the risk of tumor formation based on the methylation level of the target DMR of the test sample; (3) optionally, a tumor tissue origin evaluation module: a module for evaluating the correlation between the test sample and origin from tumor tissue based on the methylation level of the target DMR of the test sample.
[0012] In one aspect, the present application provides a system for determining a methylation variable region DMR, the system comprising a methylation variable region DMR division module: a module for determining a methylation variable region DMR based on the depth of sequencing coverage of methylation sites and / or the degree of difference in methylation levels of adjacent methylation sites.
[0013] In one aspect, the present application provides a system for evaluating the correlation between a test sample and the risk of tumor formation, the system comprising a tumor formation risk evaluation module: a module for evaluating the correlation between the test sample and the risk of tumor formation based on the methylation level of the DMR of the test sample, and the module includes a module for reducing the influence of the age factor of the subject on the evaluation result, wherein the test sample is derived from the subject.
[0014] One In one aspect, the present application provides a system for evaluating the correlation between a test sample and the origin of tumor tissue, the system comprising a tumor tissue origin evaluation module: a module for evaluating the correlation between the test sample and the origin of tumor tissue by a multi-class classification method and a logistic regression method based on the methylation level of the DMR of the test sample.
Advantages of the Invention
[0015] The present application contributes to the accurate prediction and evaluation of the risks of various cancers by providing a method with low cost and high precision.
[0016] Those skilled in the art will be able to easily understand other aspects and advantages of the present application from the following detailed description. In the following detailed description, only exemplary embodiments of the present application are shown and described. As will be recognized by those skilled in the art, the content of the present application allows those skilled in the art to make changes to the disclosed specific embodiments without departing from the spirit and scope of the present invention to which the present application pertains. Therefore, the descriptions in the accompanying drawings and the specification of the present application are merely illustrative and are not intended to be limiting.
Brief Description of the Drawings
[0017] The specific features of the invention according to the present application are shown in the appended claims. The features and advantages of the present invention related to the present application can be better understood by referring to the exemplary embodiments and the accompanying drawings described in detail below. A brief description of the accompanying drawings is set forth below.
[0018]
Figure 1
Figure 2
Figures 3A - 3C
Figure 4
Figure 5
Figures 6A - 6F
Figures 7A - 7F
Figures 8A - 8D
Figure 9
DETAILED DESCRIPTION OF THE INVENTION
[0019] Hereinafter, embodiments of the present invention will be described by specific specific embodiments, but other advantages and effects of the present invention can be easily understood by those skilled in the art from the content disclosed herein.
[0020] Definition of Terms In the present application, the terms "next-generation sequencing (NGS)", "high-throughput sequencing", or "next-generation sequencing" generally refer to next-generation high-throughput sequencing technologies and even higher-throughput sequencing methods developed thereafter. Next-generation sequencing platforms include, but are not limited to, existing sequencing platforms such as Illumina. As sequencing technologies continue to evolve, it will be understood by those skilled in the art that other sequencing methods and apparatuses may also be adopted for use in the method of the present invention. For example, next-generation sequencing may have advantages such as high sensitivity, high throughput, high sequencing depth, or low cost. Depending on the development history, influence, different sequencing principles and technologies, etc., the following main types of sequencing methods exist: Massively Parallel Signature Sequencing (MPSS), Polony Sequencing, 454 pyrosequencing, Illumina (Solexa) sequencing, Ion semiconductor sequencing, DNA nano-ball sequencing, the DNA nanoarray and probe-anchor ligation complex sequencing method of Complete Genomics, etc. The said next-generation sequencing enables a detailed and comprehensive analysis of one species of transcriptome and genome, and is thus also called deep sequencing. For example, the method of the present application can also be applied to first-generation gene sequencing, next-generation gene sequencing, third-generation gene sequencing, or single-molecule sequencing (SMS).
[0021] In the present application, the term "test sample" generally refers to a sample to be tested. For example, for one or more gene regions of a test sample, the presence or absence of a modified state can be detected.
[0022] In the present application, the terms "polynucleotide", "nucleotide", "nucleic acid" and "oligonucleotide" are used interchangeably. These refer to polymeric forms of nucleotides (deoxyribonucleotides or ribonucleotides) of any length, or analogs thereof. A polynucleotide can have any three-dimensional structure and can perform any known or unknown function. The following are non-limiting examples of polynucleotides: coding or non-coding regions of genes or gene fragments, loci (loci) identified according to linkage analysis, exons, introns, messenger RNA (mRNA), transporter RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozyme, cDNA, recombinant polynucleotide, branched polynucleotide, plasmid, vector, isolated DNA having any sequence, isolated RNA having any sequence, nucleic acid probe, primer, and linker. A polynucleotide can contain one or more modified nucleotides such as methylated nucleotides and nucleotide analogs.
[0023] In the present application, the term "methylation" generally refers to the methylation state of a gene fragment, nucleotide or its base in the present application. For example, a DNA fragment in which the gene of the present application is present may be methylated on one or more strands. For example, a DNA fragment in which the gene of the present application is present may be methylated at one site or at multiple sites.
[0024] In this application, the term "human reference genome" generally refers to the human genome that can perform a reference function in gene sequencing. Information regarding the human reference genome can be referred to UCSC. The human reference genome is available as different versions, for example, hg19, GRCH37 or ensembl 75.
[0025] In this application, the term "machine learning model" generally refers to a system or program instructions and / or an aggregate of data arranged to implement an algorithm, process or mathematical model. In this application, the algorithm, process or mathematical model can be evaluated based on a given input and provide a desired output. In this application, the parameters of the machine learning model may not be explicitly programmed, and in the conventional sense, the machine learning model may not be explicitly designed to follow specific rules to provide a desired output for a given input. For example, the use of the machine learning model can mean that the machine learning model and / or the data structure / rule set as the machine learning model are trained by a machine learning algorithm.
[0026] In this application, the term "comprising" generally means including the specifically designated features, but does not mean excluding other factors.
[0027] In this application, the term "about" generally means varying within the range of ±0.5 to 10% of the designated value, for example, within the range of ±0.5%, ±1%, ±1.5%, ±2%, ±2.5%, ±3%, ±3.5%, ±4%, ±4.5%, ±5%, ±5.5%, ±6%, ±6.5%, ±7%, ±7.5%, ±8%, ±8.5%, ±9%, ±9.5%, or ±10% of the designated value.
[0028] In order to detect six types of cancers with high incidence and high fatality rate, such as lung cancer, colorectal cancer, liver cancer, ovarian cancer, pancreatic cancer, and esophageal cancer, in this application, a new algorithm that combines the public database (TCGA) and in-house data mining to simultaneously compare the methylation variants and spatial positions of the genome is adopted, and a total of 2,536 variant regions (differentially methylated region: DMR) highly correlated with cancer are screened.
[0029] Summary of the Invention In one aspect, the present application provides a method for evaluating the correlation between a test sample and the risk of tumor formation and / or the origin of tumor tissue, comprising: (1) a step of dividing the methylation variable region DMR: determining a plurality of target DMRs for evaluation based on the depth of sequencing coverage of methylation sites and / or the degree of difference in methylation levels of adjacent methylation sites; (2) a step of evaluating the risk of tumor formation: evaluating the correlation between the test sample and the risk of tumor formation based on the methylation level of the target DMR of the test sample; (3) optionally, a step of evaluating the origin of tumor tissue: evaluating the correlation between the test sample and the origin of tumor tissue based on the methylation level of the target DMR of the test sample. For example, the method for evaluating the correlation between the test sample of the present application and the risk of tumor formation and / or the origin of tumor tissue may include: (1) determining a plurality of target DMRs for evaluation based on the depth of sequencing coverage of methylation sites and / or the degree of difference in methylation levels of adjacent methylation sites; (2) evaluating the correlation between the test sample and the risk of tumor formation based on the methylation level of the target DMR of the test sample; (3) optionally, evaluating the correlation between the test sample and the origin of tumor tissue based on the methylation level of the target DMR of the test sample.
[0030] In one aspect, the present application provides a method for determining a methylation variable region DMR, comprising the steps of dividing the methylation variable region DMR; and determining the methylation variable region DMR based on the depth of sequencing coverage of methylation sites and / or the degree of difference in methylation levels of adjacent methylation sites.
[0031] In one aspect, the present application provides a method for evaluating the correlation between a test sample and the risk of tumor formation, comprising the step of evaluating the risk of tumor formation: evaluating the correlation between the test sample and the risk of tumor formation based on the methylation level of the DMR of the test sample, including reducing the influence of the age factor of the subject on the evaluation result, wherein the test sample is derived from the subject.
[0032] In one aspect, the present application provides a method for evaluating the correlation between a test sample and tumor tissue origin, comprising the step of evaluating tumor tissue origin: evaluating the correlation between the test sample and tumor tissue origin by a multi-class classification method and a logistic regression method based on the methylation level of the DMR of the test sample.
[0033] For example, the method may include determining the DMR based on the depth of sequencing coverage of methylation sites and the degree of difference in methylation levels between a methylation site and its adjacent methylation site. For example, the degree of difference in methylation levels can refer to the difference in methylation levels. For example, the degree of difference in methylation levels can refer to the absolute value of the difference in methylation levels. For example, the present application can determine DMR regions having substantially the same methylation level based on the degree of difference in methylation levels between a methylation site and its adjacent methylation site. For example, the present application can make the division of DMR regions more accurate based on the depth of sequencing coverage of methylation sites. For example, the data information of sites with higher coverage depth is more reliable.
[0034] For example, the method may include determining an absolute value of a difference in methylation levels between a methylation site and an adjacent methylation site, and determining, based on the absolute value of the difference, whether the methylation site and the adjacent methylation site are divided into the same DMR. For example, the method may include determining a weight of the absolute value of the difference, where the weight of the absolute value of the difference is determined based on the depth of sequencing coverage of the methylation site. For example, JPEG2025524966000143.jpg16170
[0035] For example, the method may include determining an absolute value of a difference in methylation levels between a methylation site and an adjacent methylation site, and determining a weight of the absolute value of the difference to determine the degree of difference in methylation levels, where the weight of the absolute value of the difference is determined based on the depth of sequencing coverage of the methylation site.
[0036] JPEG2025524966000144.jpg31170JPEG2025524966000145.jpg16170JPEG2025524966000146.jpg38170
[0037] JPEG2025524966000147.jpg9170 If so, it is determined that the methylation site and the adjacent methylation site are divided into the same DMR.
[0038] For example, the method may further include determining the degree of variation in methylation level of the DMR based on a difference in the degree of difference in methylation levels between a methylation site at the DMR and a methylation site at an intermediate position of the DMR. For example, the intermediate position means an intermediate position in a physical position. For example, if M is odd and the DMR has M methylation sites, the intermediate position may refer to the methylation site approximately (M + 1) / 2 from upstream to downstream. For example, if M is even and the DMR has M methylation sites, the intermediate position may refer to the methylation site approximately M / 2 or M / 2 + 1 from upstream to downstream.
[0039] For example, by determining the degree of variation between the degree of methylation difference of each methylation site in the candidate DMR and the degree of methylation difference of the methylation site at the intermediate position, a more preferable DMR among the candidate DMRs is screened.
[0040] Determining DMRs of less than about 1, such as JPEG2025524966000148.jpg38170 and JPEG2025524966000149.jpg16170, is used to evaluate the correlation between the test sample and the tumor formation risk and / or the origin of the tumor tissue.
[0041] For example, the method may include, based on the methylation level of the DMR of the test sample, evaluating by a binary classification model that the test sample has a risk of tumor formation, and the evaluation method reduces the influence of the age factor of the subject on the evaluation result of the correlation between the test sample and the tumor formation risk and / or the origin of the tumor tissue, and the test sample is derived from the subject.
[0042] For example, the binary classification model may include a support vector machine SVM model. For example, the method may include introducing a penalty term based on the age factor in the SVM model. For example, the method may include introducing a penalty term based on the age factor in the SVM model by the method of Hilbert - Schmidt independence criterion. For example, all methods of introducing a penalty term of machine learning in the present application can be used to reduce the influence of the age factor in the present application.
[0043] For example, the method may include performing machine learning training on training samples with known presence or absence of tumor formation according to the following formula. JPEG2025524966000150.jpg29170
[0044] The following formula is used to determine the training parameters. JPEG2025524966000151.jpg42170JPEG2025524966000152.jpg33170JPEG2025524966000153.jpg40170JPEG2025524966000154.jpg16170
[0045] For example, the method may include determining a classification probability by a multi-class classification method based on the methylation level of the DMR of the test sample, and fitting the classification probability by logistic regression to evaluate the correlation between the test sample and the origin of the tumor tissue. For example, the method determines the classification probability by pairwise voting. For example, the method can determine the classification probability by various multi-class classification methods in the art. For example, the method fits the classification probability by multiple linear regression (MLR).
[0046] For example, the method may include performing a regression analysis on a training sample with a known tissue origin according to the following formula, JPEG2025524966000155.jpg41170JPEG2025524966000156.jpg35170JPEG2025524966000157.jpg43170
[0047] For example, the method corrects the tissue origin of the training sample based on the probability that a tumor is formed in the sample. For example, the method may include performing the correction before obtaining the classification probability by the pairwise voting. For example, the method may include performing the correction after obtaining the classification probability by the pairwise voting and before the multiple linear regression analysis. For example, the method may include performing the correction based on the maximum likelihood estimation method.
[0048] For example, the method may include performing the correction according to the following formula, JPEG2025524966000158.jpg40170
[0049] For example, by maximizing the expected value of an expression and determining weights, it is possible to correct the class derived from the tissue based on whether a tumor is formed in the sample. For example, by evaluating a sample with tumor formation, the information derived from the tissue becomes more reliable.
[0050] In one aspect, the present application provides a system for evaluating the correlation between a test sample and the risk of tumor formation and / or the origin of tumor tissue, the system comprising: (1) a methylation variable region division module: a module for determining a plurality of target DMRs to be used in the evaluation based on the depth of sequencing coverage of methylation sites and / or the degree of difference in methylation levels of adjacent methylation sites; (2) a tumor formation risk evaluation module: a module for evaluating the correlation between the test sample and the tumor formation risk based on the methylation level of the target DMR of the test sample; (3) optionally, a tumor tissue origin evaluation module: a module for evaluating the correlation between the test sample and the origin of tumor tissue based on the methylation level of the target DMR of the test sample.
[0051] In one aspect, the present application provides a system for determining a methylation variable region DMR, the system comprising a methylation variable region DMR division module: a module for determining a methylation variable region DMR based on the depth of sequencing coverage of methylation sites and / or the degree of difference in methylation levels of adjacent methylation sites.
[0052] In one aspect, the present application provides a system for evaluating the correlation between a test sample and the risk of tumor formation, the system comprising a tumor formation risk evaluation module: a module for evaluating the correlation between the test sample and the tumor formation risk based on the methylation level of the DMR of the test sample, the module including a module for reducing the influence of the age factor of the subject on the evaluation result, wherein the test sample is derived from the subject.
[0053] In one aspect, the present application provides a system for evaluating the correlation between a test sample and a tumor tissue origin, the system including a tumor tissue origin evaluation module: a module for evaluating the correlation between the test sample and the tumor tissue origin by a multi-class classification method and a logistic regression method based on the methylation level of the DMR of the test sample.
[0054] For example, the system may include determining the DMR based on the depth of sequencing coverage of the methylation site and the degree of difference in methylation levels between the methylation site and its adjacent methylation site. For example, the degree of difference in methylation levels can refer to the difference in methylation levels. For example, the degree of difference in methylation levels can refer to the absolute value of the difference in methylation levels. For example, the present application can determine a DMR region having substantially the same methylation level based on the degree of difference in methylation levels between a methylation site and its adjacent methylation site. For example, the present application can make the division of the DMR region more accurate based on the depth of sequencing coverage of the methylation site. For example, the data information of a site with a higher coverage depth is more reliable.
[0055] For example, the system may include determining the absolute value of the difference in methylation levels between a methylation site and its adjacent methylation site, and determining whether the methylation site and its adjacent methylation site are divided into the same DMR based on the absolute value of the difference. For example, the system may include determining the weight of the absolute value of the difference, and the weight of the absolute value of the difference is determined based on the depth of sequencing coverage of the methylation site. JPEG2025524966000159.jpg18170
[0056] For example, the system may include determining an absolute value of a difference in methylation levels between a methylation site and an adjacent methylation site, and determining a weight of the absolute value of the difference to determine a degree of difference in methylation levels, wherein the weight of the absolute value of the difference is determined based on a depth of sequencing coverage of the methylation site.
[0057] JPEG2025524966000160.jpg30170JPEG2025524966000161.jpg15170JPEG2025524966000162.jpg45170
[0058] JPEG2025524966000163.jpg7170 If so, it is determined that the methylation site and its adjacent methylation site are split into the same DMR.
[0059] For example, the system further includes determining a degree of variation in methylation level of the DMR based on a difference in degree of difference in methylation levels between a methylation site of the DMR and a methylation site at an intermediate position of the DMR. For example, the intermediate position means an intermediate position in a physical position. For example, when M is odd and the DMR has M methylation sites, the intermediate position may refer to the (M + 1) / 2-th methylation site from upstream to downstream. For example, when M is even and the DMR has M methylation sites, the intermediate position may refer to the M / 2-th or M / 2 + 1-th methylation site from upstream to downstream.
[0060] For example, by determining a degree of variation between a degree of difference in methylation of each methylation site in a candidate DMR and a degree of difference in methylation of a methylation site at an intermediate position, a more preferable DMR among the candidate DMRs is screened.
[0061] JPEG2025524966000164.jpg37170JPEG2025524966000165.jpg16170 Determining a DMR of less than about 1 is used to evaluate the correlation between the test sample and tumor formation risk and / or tumor tissue origin.
[0062] For example, the system may include evaluating, by a binary classification model, that a test sample has a risk of tumor formation based on the methylation level of DMR of the test sample, and the evaluation system reduces the influence of the age factor of the subject on the evaluation result of the correlation between the test sample and the risk of tumor formation and / or being derived from tumor tissue, and the test sample is derived from the subject.
[0063] For example, the binary classification model may include a support vector machine (SVM) model. For example, the system may include introducing a penalty term based on the age factor in the SVM model. For example, the system may include introducing a penalty term based on the age factor in the SVM model by a Hilbert - Schmidt independence criterion method. For example, all methods of introducing a penalty term for machine learning in the present application can be used to reduce the influence of the age factor in the present application.
[0064] For example, the system may include performing machine learning training on a training sample for which the presence or absence of tumor formation is known according to the following formula. JPEG2025524966000166.jpg29170
[0065] To determine the training parameters, the following formula is used. JPEG2025524966000167.jpg48170JPEG2025524966000168.jpg35170JPEG2025524966000169.jpg45170JPEG2025524966000170.jpg16170
[0066] For example, the system may include determining a classification probability by a multi-class classification method based on the methylation level of the DMR of the test sample, and fitting the classification probability by logistic regression to evaluate the correlation between the test sample and the origin of the tumor tissue. For example, the system determines the classification probability by pairwise voting. For example, the system can determine the classification probability by various multi-class classification method modules in the art. For example, the system fits the classification probability by multiple linear regression MLR.
[0067] For example, the system may include performing a regression analysis on a training sample with a known tissue origin according to the following formula, JPEG2025524966000171.jpg43170JPEG2025524966000172.jpg26170JPEG2025524966000173.jpg34170JPEG2025524966000174.jpg16170
[0068] For example, the system corrects the tissue origin of the training sample based on the probability that a tumor is formed in the sample. For example, the system may include performing the correction before obtaining the classification probability by the pairwise voting. For example, the system may include performing the correction after obtaining the classification probability by the pairwise voting and before the multiple linear regression analysis. For example, the system may include performing the correction based on the maximum likelihood estimation method.
[0069] For example, the system may include performing the correction according to the following formula, JPEG2025524966000175.jpg41170
[0070] For example, by maximizing the expected value of an expression and determining weights, it is possible to correct the class of tissue origin based on whether a tumor is formed in a sample. For example, by evaluating a sample with tumor formation, the information derived from the tissue becomes more reliable.
[0071] In one aspect, the present application provides a storage medium storing a program capable of executing the method described in the present application. For example, the non-volatile computer-readable storage medium may include a floppy disk, a flexible disk, a hard disk, a solid state storage (SSS) (e.g., a solid state drive (SSD)), a solid state card (SSC), a solid state module (SSM)), an enterprise flash drive, a magnetic tape, or any other non-transitory magnetic medium. Further, the non-volatile computer-readable storage medium may include a punch card, a paper tape, a photomarker sheet (or any other physical medium having a hole pattern or other optically distinguishable marking), a compact disc read-only memory (CD-ROM), a rewritable compact disc (CD-RW), a digital versatile disc (DVD), a Blu-ray disc (BD), and / or any other non-temporary optical medium.
[0072] In one aspect, the present application provides an apparatus including the storage medium described in the present application and, optionally, a processor coupled to the storage medium, the processor being arranged to implement the method described in the present application based on execution of a program stored in the storage medium.
[0073] <Example 1> For the sample, perform exemplary bisulfite-treated second-generation sequencing, and the obtained sequencing data includes the methylation level of the methylation site CpG and the depth of sequencing coverage. Optionally, noise removal is performed for the genomic methylation signal CpG and the noise region CHH / CHG sites. Next, for the "tumor" (C) group and the "normal" (N) group, calculate the p-values obtained by weighted logistic regression, where the explanatory variable of the logistic regression is a continuous variable, i.e., the methylation level of each CpG site, and the response variable is a binary output, i.e., (0,1) corresponding to C and N. Weighted logistic regression tests the distinction between C and N at each CpG site, and the null hypothesis is that the difference between C and N at that CpG site is not statistically significant. The weights are determined based on the depth of coverage of each CpG site.
[0074] Segmentation of DMR Based on the methylation level of the methylation site CpG and the depth of sequencing coverage, determine how to divide each region of the DMR. Specifically, calculate the methylation level of the methylation site CpG and the sequencing coverage depth according to the following formula: JPEG2025524966000176.jpg39170JPEG2025524966000177.jpg30170JPEG2025524966000178.jpg23170The larger the value, the higher the similarity of the methylation levels between adjacent CpG sites within the same group.
[0075] Figure 1 shows an exemplary situation (a theoretical exemplary demonstration and not intended to represent the actual sequencing situation).
[0076] For the first CpG site in this region, Sample A and Sample B each obtained coverage of 500 valid sequences, and Sample C obtained coverage of 200 valid sequences. In Sample A, the methylation level of this CpG site is 0.2. The methylation level of the second CpG site in Sample A is 0. When calculating the value of the coverage depth parameter P of the first CpG site in this group for the three samples, it is 0.617. JPEG2025524966000179.jpg33170
[0077] Figure 2 shows another exemplary situation (a theoretical exemplary demonstration and not intended to represent an actual sequencing situation).
[0078] When the above samples are replaced with A, B, and D (where Sample D obtained coverage of 400 valid sequences at the first CpG site), similarly, in Sample A, the methylation level of this CpG site is 0.2. The methylation level of the second CpG site in Sample A is 0. However, in this example, because the sequencing coverage depth of Sample D increased, for the three samples, the coverage depth of the first CpG site in the group JPEG2025524966000180.jpg21170
[0079] Therefore, by introducing the coverage depth of the CpG site by the method of the present application, the accuracy of DMR region segmentation can be significantly improved.
[0080] JPEG2025524966000181.jpg54170
[0081] Figures 3A - 3C show another exemplary situation (a theoretical exemplary demonstration and not intended to represent an actual sequencing situation). There are 10 JPEG2025524966000182.jpg11170
[0082] Among these, the calculation steps of the values within the DMR region indicated by Group A are as follows: JPEG2025524966000183.jpg74170
[0083] JPEG2025524966000184.jpg27170
[0084] The DMR regions screened by this method contain not only cancer mutation information of various cancer types but also tissue-specific characteristics, and the segmentation effect at the region boundaries is also high.
[0085] Figure 4 shows that for six types of cancers, namely lung cancer ( L ung C arcinoma:LC), colorectal cancer ( C olorectal C arcinoma:CRC), liver cancer ( L iver H epatocellular Carcinoma:LIHC), ovarian cancer ( O varian C arcinoma:OVCA), pancreatic cancer ( P ancreatic A denocarcinoma:PAAD), and esophageal cancer ( E sophageal C arcinoma:ESCA), it is shown that a tissue tracking accuracy of 98% (95% CI: 96% - 99%) can be achieved in five-fold cross-validation.
[0086] <Example 2> Cancer Evaluation (DOC) Modeling The content of ctDNA in the blood varies greatly depending on different cancer development stages and is susceptible to experimental batch effects. In addition, methylation variants are associated with factors such as age, disease, and ethnicity. If these are left unaddressed, they may affect the accuracy of the classification model as confounding variables. This application adopts a modeling method called Salmon. First, it quantifies the bias caused by confounding variables (quantification can be performed by, but is not limited to, the Hilbert-Schmidt independence criterion), incorporates a regularization term into the model to correct it, and improves the accuracy and generalization of the model.
[0087] Establishment of the algorithm JPEG2025524966000185.jpg25170JPEG2025524966000186.jpg20170
[0088] JPEG2025524966000187.jpg40170JPEG2025524966000188.jpg26170JPEG2025524966000189.jpg32170
[0089] Using the support vector machine (SVM) as the main classifier, JPEG2025524966000190.jpg28170The determination of the classification interface is determined by solving the following objective equation, JPEG2025524966000191.jpg34170For inseparable data, the soft-margin SVM introduces a penalty term for the training error, JPEG2025524966000192.jpg45170JPEG2025524966000193.jpg15170
[0090] JPEG2025524966000194.jpg21170JPEG2025524966000195.jpg38170JPEG2025524966000196.jpg14170
[0091] Figure 5 shows the control result of the weight arrangement of the interference correlation features in the Salmon-DOC model of the present application.
[0092] Each data point represents a blood sample used in the construction of the Salmon-DOC model. The horizontal axis is the Confusing Factor of the corresponding sample, and the vertical axis is the original Variable Coef before correction (Figure A) and the Variable Coef after correction (Figure B), respectively. Comparing before and after correction, it is shown that the weights of the interference correlation features are controlled in Salmon-DOC.
[0093] Retrospective cohort data In the present application, the accuracy of Salmon's binary classifier (cancer vs. non-cancer) was evaluated using retrospective clinical samples of six cancer types divided into a training set and a validation set.
[0094] Figures 6A - 6F show that the Salmon-DOC model of the present application in the tumor group model can efficiently detect different stages of six cancer types.
[0095] Figures 7A - 7F show that in the healthy group, the Salmon-DOC model of the present application overcomes the weakness of methylation false positives that increase with age and maintains balance in each age group (age on the horizontal axis, model cancer probability score on the vertical axis).
[0096] <Example 3> Tissue tracking (TOO) modeling Construction of the first layer of the TOO model The TOO model is essentially a multi-class classification problem. The calculation of the probability for each class is to vote on the pairwise results and select the result with the most votes. However, for the possible clinical applications of the tissue tracking model, simply obtaining the classification results is not sufficient. In order to enable the assembly of the model, it is necessary to generate the classification probabilities.
[0097] Therefore, the first step of the Salmon-TOO model in this application is to quantify the binary voting results. This quantification is achieved through probability calculation. JPEG2025524966000197.jpg28170JPEG2025524966000198.jpg35170JPEG2025524966000199.jpg96170
[0098] Construction of the second - layer TOO model The second layer of the Salmon - TOO model is to apply MLR to different classes.
[0099] JPEG2025524966000200.jpg42170JPEG2025524966000201.jpg83170
[0100] JPEG2025524966000202.jpg24170
[0101] JPEG2025524966000203.jpg22170JPEG2025524966000204.jpg23170
[0102] In the Salmon-DOC model, it can be seen that for certain cancer types, it is determined as negative, and for certain cancer types, it is determined as positive. Therefore, when performing tracking modeling for this determination, weight correction based on the maximum likelihood estimation method is performed for the tissue class. Taking binary logistic regression as an example, it can be interpreted as follows. JPEG2025524966000205.jpg23170
[0103] Retrospective cohort data All data of the retrospective cohort are randomly split into a training set and a validation set at a ratio of 1:1. First, the tracking evaluation results are obtained by cross-validation using the training set, and the model parameters are continuously optimized during this process and finally locked. Finally, all data of the validation set are evaluated using the locked model to obtain the tracking results. In the training set of the tracking model, the total number of samples of 6 cancer types is 300 cases, and the number of each cancer type and each stage is relatively balanced. There are 36 cases of lung cancer (the number of cases in stages I-IV is 4 / 12 / 5 / 15 respectively), 62 cases of colorectal cancer (the number of cases in stages I-IV is 8 / 18 / 18 / 18 / 18 respectively), 74 cases of liver cancer (the number of cases in stages I-IV is 25 / 14 / 22 / 13 respectively), 48 cases of ovarian cancer (the number of cases in stages I-IV is 1 / 4 / 38 / 5 respectively), 40 cases of pancreatic cancer (the number of cases in stages I-IV is 3 / 6 / 13 / 18 respectively), and 42 cases of esophageal cancer (the number of cases in stages I-IV is 5 / 10 / 15 / 12 respectively). The tracking model validation set has a total of 224 samples, including 31 cases of lung cancer (the number of cases in stages I-IV is 4 / 5 / 12 / 10 respectively), 52 cases of colorectal cancer (the number of cases in stages I-IV is 7 / 15 / 13 / 17 respectively), 55 cases of liver cancer (the number of cases in stages I-IV is 17 / 11 / 20 / 7 respectively), 27 cases of ovarian cancer (the number of cases in stages I-IV is 3 / 4 / 8 / 12 respectively), 25 cases of pancreatic cancer (the number of cases in stages I-IV is 4 / 6 / 6 / 9 respectively), and 34 cases of esophageal cancer (the number of cases in stages I-IV is 4 / 7 / 8 / 15 respectively).
[0104] Figures 8A - 8D show that the tracking accuracy of the Salmon-TOO two-layer model of the present application is better than that of the one-layer model in both cross-validation and independent validation.
[0105] Figures 8A and 8B show the tracking evaluation results by cross-validation of six types of cancer data in the training sets of six types of cancers. Among them, Figure 8A is the output result after constructing only the first-layer TOO model, and the tracking accuracy is 0.87 (260 / 300). When incorporating the quasi-optimal tracking results, the accuracy becomes 0.93 (279 / 300). Figure 8B is the output result after adding the second-layer MLR model based on the first-layer TOO model. The tracking accuracy is increased to 0.90 (270 / 300). When incorporating the quasi-optimal tracking results, the accuracy can be further increased to 0.95 (284 / 300). Similarly, Figures 8C and 8D show the tracking evaluation results of independent verification of the six types of cancer data in the above verification set. Among them, Figure 8C is the output result after constructing only the first-layer TOO model, and the tracking accuracy is 0.77 (173 / 224). When incorporating the quasi-optimal tracking results, the accuracy becomes 0.87 (194 / 224). Figure 8D is the output result after adding the second-layer MLR model based on the first-layer TOO model. The tracking accuracy is increased to 0.84 (187 / 224). When incorporating the quasi-optimal tracking results, the accuracy can be further increased to 0.89 (199 / 224).
[0106] Summarizing these results, the evaluation accuracy of the Salmon-TOO two-layer tracking model of this application is superior to that of the one-layer model in both cross-validation and independent verification of the training set.
[0107] <Example 4> DOC cancer detection model Table 1A shows 94 DMR regions used in the DOC cancer detection model. JPEG2025524966000206.jpg 229170 JPEG2025524966000207.jpg 70170
[0108] Based on 94 DOC correlation DMR regions, when evaluating 100 healthy subject samples and 318 cancer-positive samples of 6 cancer types in independent validation set 1, the overall sensitivity was 80.5% (256 / 318), and the overall specificity was 95% (95 / 100). While maintaining the specificity at the 90% level, the sensitivities in specific cancer types and stages are shown in the following table. JPEG2025524966000208.jpg186170
[0109] Subsequently, repeated tests were performed, and 50 out of 94 DOC regions were randomly selected for each test. While maintaining the specificity at the 90% (90 / 100) level, the sensitivity results of 6 cancer-positive samples in 5 repeated tests are shown in the following table. JPEG2025524966000209.jpg76170
[0110] <Example 5> TOO tissue tracking model Table 1B shows 103 DMR regions used in the TOO tissue tracking model. JPEG2025524966000210.jpg124170JPEG2025524966000211.jpg208170
[0111] Based on 103 TOO correlation DMR regions, when performing a follow-up evaluation on 473 cancer-positive samples of 6 cancer types in independent validation set 2, the initial follow-up accuracy was 63.0% (298 / 473 cases), and when incorporating the quasi-optimal follow-up results, this accuracy can be increased to 71.5% (338 / 473 cases).
[0112] Figure 9 shows the results of the obtained tissue tracking evaluation based on 103 TOO-related DMR regions.
[0113] Subsequently, 4 repeated tests were performed, 50 out of 103 TOO regions were randomly selected each time, and the follow-up accuracy results in 4 evaluations are shown in the following table. JPEG2025524966000212.jpg46170
[0114] <Example 6> Simultaneous evaluation results of DOC and TOO for 222 DMRs: Table 1C shows 222 DMR regions used in the DOC and TOO evaluation models. JPEG2025524966000213.jpg111170JPEG2025524966000214.jpg232170JPEG2025524966000215.jpg234170JPEG2025524966000216.jpg94170
[0115] In the independent validation set, when the number of markers was 222, the sensitivity and tracking accuracy at a unified specificity of 95.1% (450 / 473) were calculated for 473 negative samples and 473 positive six - cancer samples. The evaluation results of tumor examination and tissue tracking are shown in the following table. JPEG2025524966000217.jpg61170JPEG2025524966000218.jpg47170
[0116] The foregoing detailed description is provided by way of explanation and example and is not intended to limit the scope of the appended claims. Multiple variations of the embodiments currently recited in this application will be apparent to those skilled in the art and are retained within the scope of the appended claims and their equivalent embodiments.
Claims
1. A method for evaluating the correlation between a test sample and the risk of tumor formation and / or origin from tumor tissue, comprising: (1) Step of dividing the methylation variable region DMR: Based on the depth of sequencing coverage of methylation sites and / or the degree of difference in methylation levels of adjacent methylation sites, determining a plurality of target DMRs to be used in the evaluation; (2) Step of evaluating the risk of tumor formation: Based on the methylation level of the target DMR of the test sample, evaluating the correlation between the test sample and the risk of tumor formation; (3) Optionally, step of evaluating origin from tumor tissue: Based on the methylation level of the target DMR of the test sample, evaluating the correlation between the test sample and origin from tumor tissue; A method characterized by comprising the above.
2. A method for determining the methylation variable region DMR, the method comprising the step of dividing the methylation variable region DMR: Based on the depth of sequencing coverage of methylation sites and / or the degree of difference in methylation levels of adjacent methylation sites, determining the methylation variable region DMR. A method characterized by comprising the above.
3. A method for evaluating the correlation between a test sample and the risk of tumor formation, the method comprising the step of evaluating the risk of tumor formation: Based on the methylation level of the DMR of the test sample, evaluating the correlation between the test sample and the risk of tumor formation, the method further comprising the step of reducing the influence of the age factor of the subject on the evaluation result, wherein the test sample is derived from the subject. A method characterized by comprising the above.
4. A method for evaluating the correlation between a test sample and origin from tumor tissue, the method comprising the step of evaluating origin from tumor tissue: Based on the methylation level of the DMR of the test sample, evaluating the correlation between the test sample and origin from tumor tissue by a multi-class classification method and a logistic regression method. A method characterized by comprising the above.
5. The method according to any one of claims 1 to 4, comprising determining the absolute value of the difference in methylation levels between a methylation site and an adjacent methylation site, and based on the absolute value of the difference, determining whether the methylation site and the adjacent methylation site are divided into the same DMR.
6. The method according to claim 5, comprising determining a weight of the absolute value of the difference, wherein the weight of the absolute value of the difference is determined based on a depth of sequencing coverage of a methylation site.
7.
8. The method according to claim 7, wherein it is determined that the methylation site and its adjacent methylation site are divided into the same DMR.
9. The method according to any one of claims 1 to 8, further comprising determining a degree of variation in the methylation level of the DMR based on a difference in the degree of difference in the methylation levels between the methylation site of the DMR and the methylation site at the intermediate position of the DMR.
10. The method according to claim 9, which is the degree of difference in the methylation levels of the methylation sites at the intermediate positions of the MR regions.
11.
12. The method according to any one of claims 1 to 11, comprising evaluating a correlation between a test sample and a tumor formation risk by a binary classification model based on the methylation level of the DMR of the test sample, wherein the evaluation method reduces an influence of an age factor of a subject on an evaluation result of a correlation between the test sample and the tumor formation risk and / or origin from a tumor tissue, and the test sample is derived from the subject.
13. The method according to claim 12, wherein the binary classification model includes a support vector machine SVM model.
14. The method according to claim 13, comprising introducing a penalty term based on the age factor into the SVM model.
15. The method according to claim 14, comprising introducing a penalty term based on the age factor into the SVM model by a Hilbert - Schmidt independence criterion method.
16. Comprising performing machine learning training on training samples known to have tumor formation or known not to have tumor formation according to the following formula: The following formula is used to determine training parameters. The method according to any one of claims 1 to 15.
17. The method according to any one of claims 1 to 16, comprising determining a classification probability by a multi - class classification method based on the methylation level of the DMR of the test sample, and fitting the classification probability by logistic regression to evaluate a correlation between the test sample and origin from a tumor tissue.
18. The method according to claim 17, wherein the classification probability is determined by pairwise voting.
19. The method according to any one of claims 17 to 18, wherein the classification probability is applied by multiple linear regression MLR.
20. The method according to any one of claims 1 to 19, comprising performing a regression analysis on a training sample with a known tissue origin according to the following formula: The method according to any one of claims 1 to 19.
21. The method according to any one of claims 1 to 20, wherein the tissue origin of the training sample is corrected based on the probability that a tumor is formed in the sample.
22. The method according to claim 21, comprising performing the correction before the multiple linear regression analysis after obtaining the classification probability by the pairwise voting.
23. The method according to any one of claims 21 to 22, comprising performing the correction based on the maximum likelihood estimation method.
24. The method according to any one of claims 21 to 23, comprising performing the correction according to the following formula: representing the probability of tumor formation The method according to any one of claims 21 to 23.
25. A storage medium storing a program capable of executing the method according to any one of claims 1 to 24.
26. An apparatus, the apparatus comprising the storage medium according to claim 25, and optionally comprising a processor coupled to the storage medium, the processor being arranged to implement the method according to any one of claims 1 to 24 based on execution of the program recorded on the storage medium.
27. A system for evaluating the correlation between a test sample and the risk of tumor formation and / or the origin of tumor tissue, (1) A methylation variable region division module: a module for determining a plurality of target DMRs used for evaluation based on the depth of coverage of the methylation site sequencing and / or the degree of difference in the methylation level of adjacent methylation sites; (2) A tumor formation risk evaluation module: a module for evaluating the correlation between the test sample and the tumor formation risk based on the methylation level of the target DMR of the test sample; (3) Optionally, a tumor tissue origin evaluation module: a module for evaluating the correlation between the test sample and the tumor tissue origin based on the methylation level of the target DMR of the test sample. A system characterized by comprising the same.
28. A system for determining a methylation variable region DMR, comprising a methylation variable region DMR splitting module: a module for determining a methylation variable region DMR based on the depth of sequencing coverage of methylation sites and / or the degree of difference in methylation levels of adjacent methylation sites.
29. A system for evaluating the correlation between a test sample and the risk of tumor formation, comprising a tumor formation risk evaluation module: a tumor formation risk evaluation module for evaluating the correlation between a test sample and the risk of tumor formation based on the methylation level of the DMR of the test sample, the module including a module for reducing the influence of the age factor of the subject on the evaluation result, wherein the test sample is derived from the subject.
30. A system for evaluating the correlation between a test sample and the origin of tumor tissue, comprising a tumor tissue origin evaluation module: a module for evaluating the correlation between a test sample and the origin of tumor tissue by a multi-class classification method and a logistic regression method based on the methylation level of the DMR of the test sample.
Citation Information
Patent Citations
Cancer detection and classification using methylome analysis
WO2019010564A1
Methods and systems for high-depth sequencing of methylated nucleic acid
WO2020243609A1