Systems and methods for cancer screening

By acquiring and analyzing data on cfDNA methylation levels, tumor fractions, tumor ploidy, and read distribution, and using a generalized linear regression model to predict cancer probability, this approach solves the problems of high cost and insufficient accuracy in existing cancer screening technologies, achieving non-invasive, low-cost, and highly sensitive screening.

CN114974430BActive Publication Date: 2026-04-03BIOCHAIN BEIJING SCI & TECH +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-25
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing cancer screening methods are costly, invasive, and lack accuracy, making it difficult to achieve early, highly sensitive, and highly specific screening.

Method used

A system is employed, comprising a data acquisition module and a cancer calculation module. By acquiring data on the methylation level, tumor score, tumor ploidy, cfDNA fragment length characteristics, and read distribution of the target region of the subject, the system uses a generalized linear regression model to predict the cancer probability and combines it with a bisulfite treatment module to process cfDNA.

Benefits of technology

It achieves non-invasive, low-cost, highly sensitive, and highly specific cancer screening, enabling early detection of cancer and reducing the harm of invasive testing. It is suitable for asymptomatic individuals and prognostic testing of cancer patients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114974430B_ABST
    Figure CN114974430B_ABST
Patent Text Reader

Abstract

This application provides a system for cancer screening, comprising: a data acquisition module for acquiring data on methylation levels, tumor scores, tumor ploidy, cfDNA fragment length characteristics, and read distribution in a target region of a subject; and a cancer calculation module for predicting the probability of a subject developing cancer based on the data acquired in the data acquisition module. This application utilizes a very simple model and five biomarkers to provide a non-invasive cancer screening method that significantly reduces the cost and improves the accuracy of cancer screening, exhibiting very high sensitivity and specificity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of biotechnology, specifically relating to a system and method for cancer screening. Background Technology

[0002] Early cancer screening can effectively improve patient survival and incidence rates. Taking colorectal cancer as an example, according to the "2020 Chinese Guidelines for Colorectal Cancer Screening and Early Diagnosis and Treatment," the development of colorectal cancer mostly follows an "adenoma-carcinoma" sequence, generally taking 5-10 years from precancerous lesions to cancer, providing a valuable time window for early diagnosis and clinical intervention. Furthermore, the prognosis of colorectal cancer is closely related to the diagnostic stage; the 5-year relative survival rate for stage I colorectal cancer exceeds 90%, while the 5-year relative survival rate for stage IV colorectal cancer with distant metastasis is below 15%. Numerous studies and practices have shown that colorectal cancer screening and early diagnosis and treatment can effectively reduce the incidence and mortality of colorectal cancer. For example, in the United States, since around 1980, the population undergoing colonoscopy screening has gradually increased. Initially, more patients were discovered due to screening, leading to a slight increase in the incidence of colorectal cancer, but subsequently, the incidence and mortality rates of colorectal cancer in the United States have significantly decreased. For other types of cancer, early screening and diagnosis are also beneficial for early intervention and treatment, improving patient survival rates.

[0003] Cell-free DNA (cfDNA) is a small, cell-free DNA fragment found in peripheral blood, originating from normal or tumor cells and their metabolism. It contains genetic information such as somatic mutations and DNA methylation. ctDNA methylation can be detected early in tumor development and exhibits good stability. Furthermore, the length distribution and tumor abundance of cfDNA are also gaining attention in detection. Currently, DNA methylation, fragment length, and tumor heterogeneity are finding applications in early cancer screening. Summary of the Invention

[0004] In view of the problems of the prior art, the purpose of this application is to provide a system and method for cancer screening.

[0005] Specifically, the following technical solutions are involved:

[0006] 1. A system for cancer screening, comprising:

[0007] The data acquisition module is used to acquire data on methylation levels, tumor scores, tumor ploidy, cfDNA fragment length characteristics, and read distribution in the target region of the subject; and

[0008] The cancer calculation module predicts the probability of a subject developing cancer based on data obtained from the data acquisition module, including methylation level, tumor score, tumor ploidy, cfDNA fragment length characteristics, and read distribution.

[0009] 2. The system according to claim 1, wherein,

[0010] The data acquisition module includes a sequencing module, a methylation level analysis module, a tumor score analysis module, a tumor ploidy analysis module, a cfDNA fragment length feature value analysis module, and a reads distribution analysis module.

[0011] The sequencing module is used to perform whole-genome sequencing on the subject's cfDNA;

[0012] The methylation level analysis module analyzes the methylation level of the target region based on sequencing data obtained from the sequencing module.

[0013] The tumor score analysis module analyzes tumor scores based on sequencing data obtained from the sequencing module;

[0014] The tumor ploidy analysis module analyzes tumor ploidy based on sequencing data obtained from the sequencing module;

[0015] The cfDNA fragment length feature value analysis module calculates a risk feature value for tumors related to the length of the cfDNA fragment extracted from the sequencing data obtained from the sequencing module.

[0016] The reads distribution analysis module analyzes the reads distribution from the sequencing data obtained from the sequencing module and then predicts the source of each read.

[0017] 3. The system according to item 1 or 2, wherein,

[0018] The target region includes any one or more of the following regions:

[0019] Chromosome 2, positions 223721500-223726500.

[0020] Chromosome 6, positions 170147000-170152000.

[0021] Chromosome 8, positions 182500-187500,

[0022] Chromosome 8, positions 64081500-64086500.

[0023] Chromosome 9, positions 14688500-14693500,

[0024] Chromosome 10, positions 119803500-119808500, or

[0025] Chromosome 15, positions 56285500-56290500.

[0026] 4. The system according to any one of items 1 to 3, wherein,

[0027] The methylation level of the target region is calculated based on the methylation level of each CG site in the target region, wherein the methylation level of the CG site is the ratio of the methylated cytosine detected at that site to the sum of the unmethylated cytosine and the unmethylated cytosine in all detected sequence results of that site;

[0028] Tumor fraction refers to the proportion of ctDNA in cfDNA, that is, the proportion of cell-free DNA released by tumor cells in the total cfDNA;

[0029] Tumor ploidy refers to the variation caused by abnormalities in chromosome structure and number, that is, the number of ploids in a tumor;

[0030] The cfDNA fragment length feature value analysis module calculates the risk feature value of tumor related to the length of the cfDNA fragment extracted from the sequencing data obtained from the sequencing module. This is based on the ratio between short and long fragments in the cfDNA fragment length, and the cfDNA fragment length feature value is calculated through a gradient boosting tree model.

[0031] Preferably, the long segment is a segment with a length of 201~320bp;

[0032] Short segments are segments with a length of 150~200bp;

[0033] The reads distribution refers to the probability of a sample origin obtained by statistically analyzing each read in the cfDNA sequencing data, that is, the probability that a given read comes from one of three source components: healthy plasma, normal tissue, and tumor tissue.

[0034] 5. The system according to any one of items 1 to 4, wherein,

[0035] cfDNA sequencing data is cfDNA sequencing data after removing low-quality sequencing fragments;

[0036] Preferably, the cfDNA sequencing data is the sequencing data after removing low-quality sequencing fragments and further excluding sequencing data in the low alignment rate range.

[0037] 6. The system according to any one of items 1 to 5, wherein,

[0038] The cancer calculation module pre-stores a formula fitted from known sample data, including methylation levels, tumor fractions, tumor ploidy, cfDNA fragment length characteristics, and read distribution, to predict the probability of a subject developing cancer. This formula is Equation 1.

[0039] Formula 1

[0040] P represents the probability that the subject will develop cancer;

[0041] X1 represents the tumor score;

[0042] X2 represents the ploidy of the tumor;

[0043] X3 represents the methylation level;

[0044] X4 is the cfDNA fragment length feature value;

[0045] X5 represents the reads distribution;

[0046] β0 is any value selected from 100 to 150, preferably 114.12;

[0047] β1 is any value selected from -700 to 0, preferably -686.25;

[0048] β2 is any value selected from 20 to 40, preferably 31.16; β4 is any value selected from -20 to 0, preferably -17.45;

[0049] 7. The system according to item 6, wherein β3X3 is obtained from equation two:

[0050] β3X3= a* X 31 + b* X 32+ c* X 33 + d* X 34 + e* X 35 + f* X 36 + g* X 37 Formula 2

[0051] X 31 The methylation level at positions 223721500-223726500 on chromosome 2 in the target region;

[0052] X 32 The methylation level at positions 170147000-170152000 on chromosome 6 in the target region;

[0053] X 33 The methylation level of chromosome 8 from position 182500 to 187500 in the target region;

[0054] X 34 The methylation level at positions 64081500-64086500 on chromosome 8 in the target region;

[0055] X35 The methylation level at positions 14688500-14693500 on chromosome 9 in the target region;

[0056] X 36 The methylation level of chromosome 10 from positions 119803500 to 119808500 in the target region;

[0057] X 37 The methylation level at positions 56285500-56290500 on chromosome 15 in the target region;

[0058] a can be any value selected from -300 to 300, preferably -260.66;

[0059] b is any value selected from -200 to 200, preferably -181.95;

[0060] c is any value selected from -100 to 100, preferably -85.21;

[0061] d is any value selected from -350 to 350, preferably -305.36;

[0062] e can be any value selected from -250 to 250, preferably 218.80;

[0063] f is any value selected from -100 to 100, preferably 60.85;

[0064] g is any value selected from -250 to 250, preferably -209.25.

[0065] 8. The system according to item 6, wherein β5X5 is obtained from equation 3:

[0066] β5X5=l* X 51 + m* X 52+ n* X 53 Formula 3

[0067] X 51 Distribution of reads derived from healthy plasma;

[0068] X 52 The distribution of reads originating from normal tissues;

[0069] X 53 Distribution of reads derived from tumor tissue;

[0070] l can be any value selected from 50 to 200, preferably 166.02;

[0071] m can be any value selected from 0 to 10, preferably 8.41;

[0072] n is any value selected from -0.1 to 1, preferably 0.00000001.

[0073] 9. The system according to item 1, wherein,

[0074] The system also includes a bisulfite treatment module for treating the subject's cfDNA with bisulfite.

[0075] 10. A method for cancer screening using any one of the systems described in items 1-9, the method comprising:

[0076] The sample collection process involves acquiring data on methylation levels, tumor scores, tumor ploidy, cfDNA fragment length characteristics, and read distribution in the target region of the subject.

[0077] The cancer calculation step uses data on methylation levels, tumor scores, tumor ploidy, cfDNA fragment length, and read distribution obtained from the data acquisition module to predict the probability of a subject developing cancer.

[0078] 11. The cancer screening method according to item 10, the method comprising:

[0079] The data acquisition steps include sequencing, methylation level analysis, tumor fraction analysis, tumor ploidy analysis, cfDNA fragment length feature value analysis, and reads distribution analysis.

[0080] The sequencing step is used to perform whole-genome sequencing of the subject's cfDNA;

[0081] The methylation level analysis step is based on sequencing data obtained from the sequencing module to analyze the methylation level of the target region.

[0082] The tumor score analysis step is based on analyzing tumor scores using sequencing data obtained from the sequencing module;

[0083] The tumor ploidy analysis step is based on analyzing tumor ploidy using sequencing data obtained from the sequencing module;

[0084] The cfDNA fragment length feature value analysis step calculates a length-related risk feature value for tumors based on the length of the cfDNA fragment extracted from the sequencing data obtained from the sequencing module.

[0085] The reads distribution analysis module analyzes the reads distribution from the sequencing data obtained from the sequencing steps to predict the source of each read.

[0086] The effects of the invention

[0087] This application utilizes a comprehensive set of indicators related to methylation level, tumor fraction, tumor ploidy, cfDNA fragment size, and read distribution to construct a non-invasive cancer screening method that can significantly reduce the cost of cancer screening and improve screening accuracy, while exhibiting very high sensitivity and specificity.

[0088] This model can: 1) be used in a non-invasive manner for early screening of asymptomatic individuals and prognostic testing of cancer patients, reducing the harm caused by invasive testing; and 2) have higher sensitivity and accuracy, enabling real-time monitoring.

[0089] This application presents a non-invasive cancer screening method using five biomarkers that significantly reduces the cost and improves the accuracy of cancer screening, exhibiting very high sensitivity and specificity. Furthermore, when using these five biomarkers, a relatively simple and robust generalized linear model can achieve good predictions. Generalized linear models are a crucial type of data model today, a generalization of classical linear models, widely used in modern society, particularly in the statistical analysis of data in medicine, biology, and economics. They are applicable to both discrete and continuous data. Generalized linear models greatly extend the standard linear model by fitting a function (not the conditional mean of the response variable) to the conditional mean of the response variable and assuming that the response variable follows a distribution within the exponential family (not limited to the normal distribution). The derivation of model parameter estimation is based on maximum likelihood estimation. Logistic regression is a very important model when the dependent variable is binary or multi-category. Because logistic regression does not require normality or homogeneity of variance in the data, nor does it impose requirements on the type of independent variables, it is widely used in many fields. At the same time, generalized linear models are generally trained very quickly, and the training process is also easy to understand. Attached Figure Description

[0090] Figure 1 The ROC curve of the MODE model in the test set in Example 2;

[0091] Figure 2 The ROC curves in the test set for the generalized linear regression model constructed based on tumor fraction in Example 3;

[0092] Figure 3 The ROC curves in the test set of the generalized linear regression model constructed based on tumor ploidy in Example 3;

[0093] Figure 4 The ROC curves in the test set of the generalized linear regression model constructed based on the reads distribution in Example 3;

[0094] Figure 5 The ROC curve of the generalized linear regression model constructed based on the fragment size feature value of cfDNA in Example 3 is shown in the test set.

[0095] Figure 6 The ROC curve of the generalized linear regression model based on methylation level constructed in Example 3 is shown in the test set.

[0096] Figure 7 The ROC curve of the generalized linear regression model constructed based on a combination of tumor ploidy and reads distribution in Example 4 is shown in the test set.

[0097] Figure 8 The ROC curves of a generalized linear regression model constructed based on a combination of tumor ploidy, tumor fraction, and cfDNA fragment size features were obtained on the test set. Detailed Implementation

[0098] Specific embodiments of the present application will now be described in more detail with reference to the accompanying drawings. While specific embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.

[0099] It should be noted that certain terms are used in the specification and claims to refer to specific components. Those skilled in the art will understand that different terms may be used to refer to the same component. This specification and claims do not distinguish components based on differences in terminology, but rather on differences in function. The terms "comprising" or "including" used throughout the specification and claims are open-ended and should be interpreted as "comprising but not limited to." The following descriptions in the specification are preferred embodiments for carrying out this application; however, these descriptions are for the purpose of understanding the general principles of the specification and are not intended to limit the scope of this application. The scope of protection of this application shall be determined by the appended claims.

[0100] definition

[0101] Unless specifically defined elsewhere in this document, all other technical and scientific terms used herein have the meanings commonly understood by one of ordinary skill in the art to which this application pertains.

[0102] methylation

[0103] Methylation is an important modification of proteins and nucleic acids, regulating gene expression and shutdown. It is closely related to many diseases such as cancer, aging, and Alzheimer's disease, and is one of the important research topics in epigenetics. The most common methylation modifications are DNA methylation and histone methylation.

[0104] DNA methylation refers to the methylation process occurring at the 5th carbon atom of cytosine in CpG dinucleotides. As a stable modification state, it can be inherited by newly generated daughter DNA during DNA replication under the action of DNA methyltransferases, making it an important epigenetic mechanism. During DNA methylation, methylation of the gene promoter region can lead to transcriptional silencing of tumor suppressor genes, thus it is closely related to tumorigenesis. Abnormal methylation includes hypermethylation of tumor suppressor genes and DNA repair genes, hypomethylation of repetitive DNA sequences, and loss of imprinting of certain genes, all of which are associated with the development of various tumors.

[0105] In this paper, the ROC curve can reflect the classification performance of a classifier to some extent. AUC is essentially the area under the ROC curve. AUC intuitively reflects the classification ability expressed by the ROC curve.

[0106] Whole genome methylation sequencing

[0107] Whole-genome bisulfite sequencing (WGBS) is considered the "gold standard" for methylation sequencing. Its principle involves treating the genome with bisulfite to convert unmethylated C bases to U, which are then amplified by PCR to become T, distinguishing them from the originally methylated C bases. This is then combined with high-throughput sequencing technology and alignment with a reference sequence to determine whether CpG / CHG / CHH sites are methylated.

[0108] This application provides a system for cancer screening, comprising a data acquisition module and a cancer calculation module. The data acquisition module acquires data on methylation levels, tumor fraction, tumor ploidy, cfDNA fragment size characteristics, and read distribution in a target region of the subject. The cancer calculation module predicts the probability of the subject developing cancer based on the data acquired in the data acquisition module.

[0109] In some embodiments of this application, the data acquisition module includes a sequencing module, a methylation level analysis module, a tumor score analysis module, a tumor ploidy analysis module, a cfDNA fragment length feature value analysis module, and a reads distribution analysis module;

[0110] The sequencing module is used to perform whole-genome sequencing on the subject's cfDNA;

[0111] The methylation level analysis module analyzes the methylation level of the target region based on sequencing data obtained from the sequencing module.

[0112] The tumor score analysis module analyzes tumor scores based on sequencing data obtained from the sequencing module;

[0113] The tumor ploidy analysis module analyzes tumor ploidy based on sequencing data obtained from the sequencing module;

[0114] The cfDNA fragment length feature value analysis module calculates a risk feature value for tumors related to the length of the cfDNA fragment extracted from the sequencing data obtained from the sequencing module.

[0115] The reads distribution analysis module analyzes the reads distribution from the sequencing data obtained from the sequencing module and then predicts the source of each read.

[0116] In this paper, the target region of the subject can be a specific region on the subject's chromosome. For example, the target region includes any one or more of the following regions:

[0117] Chromosome 2, positions 223721500-223726500.

[0118] Chromosome 6, positions 170147000-170152000.

[0119] Chromosome 8, positions 182500-187500,

[0120] Chromosome 8, positions 64081500-64086500.

[0121] Chromosome 9, positions 14688500-14693500,

[0122] Chromosome 10, positions 119803500-119808500, or

[0123] Chromosome 15, positions 56285500-56290500.

[0124] In one specific implementation, the target region is positions 223721500-223726500 on chromosome 2.

[0125] In one specific implementation, the target region is chromosome 6, positions 170147000-170152000.

[0126] In one specific implementation, the target region is positions 182500-187500 on chromosome 8.

[0127] In one specific implementation, the target region is chromosome 8, positions 64081500-64086500.

[0128] In one specific implementation, the target region is chromosome 9, positions 14688500-14693500.

[0129] In one specific implementation, the target region is chromosome 10, positions 119803500-119808500.

[0130] In one specific implementation, the target region is chromosome 15, positions 56285500-56290500.

[0131] The selection of target regions first utilizes the sliding window method to slide windows on the reference genome and calculate the overall methylation level of CpG sites in each window interval. For each sample, the methylation level of the corresponding window is statistically analyzed. Differentially methylated windows are identified according to the different sample groups. Window with a methylation difference greater than 0.2 is then selected as the target region, thus obtaining the aforementioned target regions.

[0132] The methylation level of the target region is calculated based on the methylation level of each CG site in the target region, wherein the methylation level of the CG site is the ratio of the methylated cytosine detected at that site to the sum of unmethylated and non-methylated cytosine in all detected sequence results of that site.

[0133] For each window, the number of CG sites within that window is counted. Since the methylation depth of cytosine at each CG site and the total depth of the site are known, the methylation level of the entire window can be calculated. This is the ratio of the sum of the methylation depths of cytosine at all CG sites to the sum of the total depths of all CG sites. Each window will yield a corresponding methylation level using this calculation method. The methylation depth of cytosine at each CG site represents the number of reads indicating methylated cytosine at that site in the sequencing results, i.e., the number of reads indicating C (cytosine) at that site. The total depth of the site represents the total number of all sequencing reads covering that site, i.e., the total number of reads indicating C or T (thymine) at that site. The methylation depth of cytosine and the total depth of the site can be directly provided after analysis by the sequencing software.

[0134] Tumor fraction, also known as tumor proportion, refers to the percentage of ctDNA in cell-free DNA (cfDNA), specifically the proportion of cfDNA from tumor cells within the total cfDNA. Peripheral circulating cell-free DNA (cfDNA) consists of DNA fragments released into the blood plasma after cell apoptosis, necrosis, or other abnormal processes. cfDNA can be used to describe various forms of cell-free DNA in peripheral blood circulation. ctDNA, derived from tumor cells, is a type of cfDNA. Therefore, tumor fraction refers to the proportion of ctDNA within cfDNA, specifically the percentage of cfDNA from tumor cells in the total cfDNA.

[0135] Among them, cfDNA refers to various forms of cell-free DNA fragments in peripheral circulating blood, and ctDNA refers to cell-free DNA derived from tumor cells in the above-mentioned cfDNA.

[0136] ctDNA was analyzed using the iChorCNA software. By performing whole-genome sequencing on cfDNA, somatic copy number alterations (SCNA) in cfDNA can be detected and tumor fractions can be quantitatively calculated.

[0137] Tumor ploidy refers to variations caused by abnormalities in chromosome structure and number, indicating the number of ploids in a tumor. Normal cells are generally diploid, but variations can cause some tumor cells to become polyploid. The presence of aneuploidy in DNA is related to the rate of carcinogenesis and the degree of precancerous lesion proliferation, and is an important indicator of the development of precancerous lesions into cancer. Tumor ploidy is determined using ichorCNA software. By performing whole-genome sequencing of cfDNA, somatic copy number alterations (SCNA) in cfDNA can be detected and tumor ploidy can be quantitatively calculated.

[0138] Studies have shown that when tumors occur, cells proliferate malignantly, and copy number variations generally occur; during tumor development, tumor cell cfDNA is released into the bloodstream. The ichoCNA software, based on in-depth evaluation of data reads, calculates copy number variation using the Hidden Markov Model (HMM) and then uses the EM algorithm to calculate tumor fraction and tumor ploidy based on the copy number variation.

[0139] cfDNA fragment size refers to the average size of a given fragment obtained from cfDNA sequencing data, i.e., the ratio of the sum of the sizes of all fragments obtained from the subject's cfDNA sequencing data to the total number of fragments. Generally, the extracted fragment information is laid out into adjacent, non-overlapping 100kb intervals based on the autosomes of the hg19 reference genome. Short fragments are defined as short (150-200 bp), long fragments as long (201-320 bp), and cov fragments are short + long. The number of short, long, and cov fragments in each interval is counted, representing fragment coverage. The LOWESS algorithm is used to correct for short, long, and cov fragments, removing coverage bias caused by GC offset. The 100kb intervals are sequentially merged into 5MB intervals, resulting in 499 non-overlapping intervals. The number of short, long, and cov fragments in each 5MB interval is counted. The short / long ratio represents the short-to-long fragment ratio. Through analysis, cov (499 intervals corresponding to 499 cov features) was selected as the feature, PCA was used for dimensionality reduction, and then a gradient boosting tree model was used to obtain the disease risk value.

[0140] The cfDNA fragment length feature value analysis module calculates the risk feature value of tumor related to the length of the cfDNA fragment extracted from the sequencing data obtained from the sequencing module. This is based on the ratio between short and long fragments in the cfDNA fragment length, and the cfDNA fragment length feature value is calculated through a gradient boosting tree model.

[0141] Preferably, the long segment is a segment with a length of 201~320bp;

[0142] Short segments are segments with a length of 150~200bp;

[0143] Read distribution refers to the probability of a cfDNA sample originating from each read in the cfDNA sequencing data, obtained by statistically analyzing each read. The CancerDetecotr software is used to predict the probability that a read originates from tumor cells based on the distribution of reads from existing cancer patients. This method employs a mixed model assumption of three tissue sources: healthy plasma, normal tissue, and tumor tissue, rather than the traditional two-component mixed assumption of healthy plasma and tumor tissue. Theoretically, the mixed model assumption can not only determine whether the sample is healthy or diseased, but also accurately distinguish whether a sample identified as diseased is cancerous. For each read, the methylation level distribution pattern characteristic value in the three groups of samples (healthy plasma, normal tissue, and tumor tissue) is calculated based on the methylation level data, thereby obtaining the tumor prediction probability of the sequenced sample. In other words, it represents the probability that a given read originates from one of the three source components: healthy plasma, normal tissue, or tumor tissue.

[0144] In one specific implementation, the cfDNA sequencing data is cfDNA sequencing data after removing low-quality sequencing fragments; preferably, the cfDNA sequencing data is sequencing data after further excluding sequencing data in low alignment rate ranges after removing low-quality sequencing fragments. Low-quality sequencing fragments are generally reads with relatively low sequencing quality, such as those aligned to multiple chromosomal locations, containing adapter data, or with an overall sequencing data size less than Q20. Removing low-quality sequencing data can significantly improve the accuracy of subsequent analysis results and reduce false positives.

[0145] The cancer calculation module of this application pre-stores formulas that are fitted based on data from known samples, including methylation levels, tumor fraction, tumor ploidy, cfDNA fragment size characteristics, and read distribution, to predict the probability of a subject developing cancer. By substituting the methylation levels, tumor fraction, tumor ploidy, cfDNA fragment size characteristics, and read distribution of the target region of the subject obtained from the data acquisition module into the formula in the cancer calculation module, the probability of the subject developing cancer can be obtained.

[0146] The formula is derived by logically fitting data on known sample methylation levels, tumor fraction, tumor ploidy, cfDNA fragment size, and read distribution using a generalized linear regression (GLM) algorithm. Specifically, the values ​​of the five biomarkers are imported into R software, integrated using GLM, with the response variable following a binomial distribution and a logistic regression link function. The regression coefficients are solved through 100 iterations based on the maximum likelihood estimation principle, ultimately fitting a MODE model.

[0147] Among them, the MODE model, i.e., Equation 1

[0148]

[0149] Formula 1

[0150] Model Interpretation: P The probability of an outcome when exposed to a certain state. logit ( P () is a variable transformation method, representing the transformation of... P conduct logit Transformation, β These are partial regression coefficients, representing the coefficients for each unit change in independent variables, assuming other independent variables remain constant. logit ( P The estimated value of ).

[0151] P represents the probability that the subject will develop cancer;

[0152] X1 represents the tumor score;

[0153] X2 represents the ploidy of the tumor;

[0154] X3 represents the methylation level;

[0155] X4 is the cfDNA fragment length feature value;

[0156] X5 represents the reads distribution;

[0157] β0 is any value selected from 100 to 150, preferably 114.12;

[0158] β1 is any value selected from -700 to 0, preferably -686.25;

[0159] β2 is any value selected from 20 to 40, preferably 31.16;

[0160] β4 is any value selected from -20 to 0, preferably -17.45;

[0161] The prediction function in R can be used to obtain the probability value predicted by MODE, thereby determining the probability that a sample has cancer.

[0162] Generalized linear models greatly extend standard linear models by fitting a function (not the conditional mean of the response variable) to the conditional mean of the response variable and assuming that the response variable follows a distribution within the exponential family (not limited to the normal distribution). The derivation of model parameter estimation is based on maximum likelihood estimation. Logistic regression is a crucial model when the dependent variable is binary or multi-category. Because logistic regression does not require normality or homogeneity of variance in the data, nor does it impose restrictions on the type of independent variables, it is widely used in numerous fields.

[0163] In some embodiments of this application, β3X3 is obtained from Equation 2:

[0164] β3X3= a* X 31 + b* X 32+ c* X 33 + d* X 34 + e* X 35 + f* X 36 + g* X 37 Formula 2

[0165] X 31 The methylation level at positions 223721500-223726500 on chromosome 2 in the target region;

[0166] X 32 The methylation level at positions 170147000-170152000 on chromosome 6 in the target region;

[0167] X 33The methylation level of chromosome 8 from position 182500 to 187500 in the target region;

[0168] X 34 The methylation level at positions 64081500-64086500 on chromosome 8 in the target region;

[0169] X 35 The methylation level at positions 14688500-14693500 on chromosome 9 in the target region;

[0170] X 36 The methylation level of chromosome 10 from positions 119803500 to 119808500 in the target region;

[0171] X 37 The methylation level at positions 56285500-56290500 on chromosome 15 in the target region;

[0172] a can be any value selected from -300 to 300, preferably -260.66;

[0173] b is any value selected from -200 to 200, preferably -181.95;

[0174] c is any value selected from -100 to 100, preferably -85.21;

[0175] d is any value selected from -350 to 350, preferably -305.36;

[0176] e can be any value selected from -250 to 250, preferably 218.80;

[0177] f is any value selected from -100 to 100, preferably 60.85;

[0178] g is any value selected from -250 to 250, preferably -209.25.

[0179] In some embodiments of this application, β5X5 is obtained from Equation 3:

[0180] β5X5=l* X 51 + m* X 52+ n* X 53 Formula 3

[0181] X 51 Distribution of reads derived from healthy plasma;

[0182] X 52 The distribution of reads originating from normal tissues;

[0183] X 53 Distribution of reads derived from tumor tissue;

[0184] l can be any value selected from 50 to 200, preferably 166.02;

[0185] m can be any value selected from 0 to 10, preferably 8.41;

[0186] n is any value selected from -0.1 to 1, preferably 0.00000001.

[0187] In one specific embodiment of this application, MODE = 114.12 + SCORE_TF + SCORE_TP + SCORE_DMV + SCORE_FS + SCORE_READS

[0188] SCORE_TF = -686.25 * Tumor Score

[0189] SCORE_TP = 31.16 * Tumor ploidy

[0190] SCORE_DMV = (-260.66 * methylation level 1) + (-181.95 * methylation level 2) + (-85.21 * methylation level 3) + (-305.36 * methylation level 4) + (218.80 * methylation level 5) + (60.85 * methylation level 61) + (-209.25 * methylation level 7)

[0191]

[0192] SCORE_FS = -17.45 * cfDNA fragment length feature value

[0193] SCORE_READS = (reads distribution of healthy plasma * 166.02) + (8.41 * reads distribution of normal tissue) + (0.00000001 * reads distribution of tumor tissue)

[0194] The system described in this application may further include a bisulfite treatment module for treating the subject's cfDNA with bisulfite. The bisulfite-treated cfDNA is then used for subsequent cfDNA sequencing.

[0195] This application also provides a method for cancer screening using the above-described system, the method comprising:

[0196] The sample collection process involves acquiring data on methylation levels, tumor fraction, tumor ploidy, cfDNA fragment size characteristics, and read distribution in the target region of the subject.

[0197] The cancer calculation step uses data on methylation level, tumor fraction, tumor ploidy, cfDNA fragment size, and read distribution obtained from the data acquisition module to predict the probability of a subject developing cancer.

[0198] In some specific implementations, the data acquisition steps include sequencing, methylation level analysis, tumor fraction analysis, tumor ploidy analysis, cfDNA fragment length feature value analysis, and reads distribution analysis.

[0199] The sequencing step is used to perform whole-genome sequencing of the subject's cfDNA;

[0200] The methylation level analysis step is based on sequencing data obtained from the sequencing module to analyze the methylation level of the target region.

[0201] The tumor score analysis step is based on analyzing tumor scores using sequencing data obtained from the sequencing module;

[0202] The tumor ploidy analysis step is based on analyzing tumor ploidy using sequencing data obtained from the sequencing module;

[0203] The cfDNA fragment length feature value analysis step calculates a length-related risk feature value for tumors based on the length of the cfDNA fragment extracted from the sequencing data obtained from the sequencing module.

[0204] The reads distribution analysis module analyzes the reads distribution from the sequencing data obtained from the sequencing steps to predict the source of each read.

[0205] Example

[0206] Example 1: Calculation of differentially methylated regions and fragment group characteristics

[0207] 1.1 cfDNA extraction and purification

[0208] 1.1.1 Plasma sample preparation:

[0209] Centrifuge the blood sample at 2000 g for 10 min at 4℃, and transfer the plasma to a new centrifuge tube. Centrifuge the plasma sample at 16000 g for 10 min at 4℃. Proceed to the next step depending on the type of collection tube used; in this experiment, an "other" type of collection tube was used.

[0210]

[0211] 1.1.2 Fracturing and Binding

[0212] 1.1.2.1. Prepare the Binding Solution / Beads Mix according to the table below, and then mix thoroughly.

[0213]

[0214] Add an appropriate volume of plasma sample.

[0215] 1.1.2.2. Thoroughly mix the plasma sample and Binding Solution / Beads Mix.

[0216] 1.1.2.3. Mix thoroughly on a rotary mixer for 10 minutes to allow the cfDNA to bind to the magnetic beads.

[0217] 1.1.2.4. Place the binding tube on the magnetic rack for 5 minutes, until the solution becomes clear and the magnetic beads are completely adsorbed on the magnetic rack.

[0218] 1.1.2.5. Carefully discard the supernatant with a pipette, and continue to keep the tube on the magnetic rack for a few minutes. Remove any remaining supernatant with a pipette.

[0219] 1.1.3 Washing

[0220] 1.1.3.1. Resuspend the beads in 1 ml of Wash Solution.

[0221] 1.1.3.2. Transfer the resuspension to a new, non-adsorbed 1.5 ml centrifuge tube. Retain the binding tube.

[0222] 1.1.3.3. Place the centrifuge tube containing the bead resuspension on a magnetic rack for 20 seconds.

[0223] 1.1.3.4. Aspirate the supernatant obtained from the separation and wash the binding tube. Collect the residual beads after washing back into the resuspension and discard the lysis / binding tube.

[0224] 1.1.3.5. Place the tube on the magnetic rack for 2 min until the solution becomes clear and the beads gather on the magnetic rack. Remove the supernatant with a 1 ml pipette.

[0225] 1.1.3.6. Leave the tube on the magnetic rack and remove as much residual liquid as possible using a 200 μL pipette.

[0226] 1.1.3.7. Remove the tube from the magnetic holder, add 1 ml of Wash Solution, and vortex for 30 seconds.

[0227] 1.1.3.8. Place on a magnetic rack for 2 min until the solution is clear and the beads gather on the magnetic rack. Remove the supernatant with a 1 ml pipette.

[0228] 1.1.3.9. Leave the tube on the magnetic rack and use a 200 μL pipette to completely remove any remaining liquid.

[0229] 1.1.3.10. Remove the tube from the magnetic rack, add 1 ml of 80% ethanol, and vortex for 30 seconds.

[0230] 1.1.3.11. Place on a magnetic rack for 2 minutes until the solution becomes clear, then remove the supernatant with a 1 ml pipette.

[0231] 1.1.3.12. Leave the tube on the magnetic rack and remove any remaining liquid using a 200 μL pipette.

[0232] 1.1.3.13. Repeat steps 10-12 above once with 80% ethanol to remove as much supernatant as possible.

[0233] 1.1.3.14. Leave the tube on the magnetic rack and allow the beads to dry in the air for 3-5 minutes.

[0234] 1.1.4 Elution of cfDNA

[0235] 1.1.4.1. Add the Elution Solution according to the table below.

[0236]

[0237] 1.1.4.2. Place on a magnetic rack for 2 minutes until the solution becomes clear, then aspirate the cfDNA from the supernatant.

[0238] 1.1.4.3. The purified cfDNA can be used immediately, or the supernatant can be transferred to a new centrifuge tube and stored at -20°C.

[0239] 1.2g DNA fragmentation and purification:

[0240] 1.2.1. According to the Qubit concentration, take 2 μg of DNA, add water to make up to 125 μl, add it to a 130 μl Covaris fragmentation tube, and set the program: 50W, 20%, 200 cycles, 250s.

[0241] 1.2.2 After the fragmentation is completed, take 1 μl of sample and use Agilent 2100 to detect the fragment. After normal fragmentation, the main peak of the sample is about 150bp-200bp.

[0242] For cfDNA samples, Agilent 2100 was used for fragment detection, and the Qubit was directly used for subsequent experiments.

[0243] 1.3 End repair, 3' end with "A":

[0244] 1.3.1. Transfer the fragmented gDNA or cfDNA from Xng into a PCR tube, add nuclease-free water to a final volume of 50 μl, add the following reagents, and vortex to mix:

[0245]

[0246] 1.3.2. Set the following program to perform the reaction on the PCR instrument:

[0247] The temperature of the hot cap is 85℃.

[0248]

[0249] 1.4 Connector connection and purification:

[0250] 1.4.1. Refer to the table below to dilute the connector to a suitable concentration in advance:

[0251]

[0252] 1.4.2. Prepare the following reagents according to the table below, gently pipette and mix well, then briefly centrifuge:

[0253]

[0254] 1.4.3. Set the following program to perform the reaction on the PCR instrument:

[0255] No heated cap.

[0256]

[0257] 1.4.4. Add purified magnetic beads to the following system for the experiment (AgencourtAMPure XP magnetic beads should be brought to room temperature and mixed thoroughly beforehand):

[0258]

[0259] 1.4.4.1. Gently whisk and mix 6 times.

[0260] 1.4.4.2. Incubate at room temperature for 5-15 minutes, then place the PCR tube on a magnetic rack for 3 minutes to allow the solution to clarify.

[0261] 1.4.4.3. Remove the supernatant, keep the PCR tube on the magnetic rack, add 200 μl of 80% ethanol solution to the PCR tube, and let it stand for 30 seconds.

[0262] 1.4.4.4. Remove the supernatant, then add 200 μl of 80% ethanol solution to the PCR tube, let it stand for 30 seconds, and then completely remove the supernatant (it is recommended to use a 10 μl pipette to remove any residual ethanol solution at the bottom).

[0263] 1.4.4.5. Let stand at room temperature for 3-5 minutes to allow the residual ethanol to evaporate completely.

[0264] 1.4.4.6. Add 22 μl of Nuclease-free water, remove the PCR tube from the magnetic rack, gently aspirate and resuspend the magnetic beads to avoid generating air bubbles, and let stand at room temperature for 2 minutes.

[0265] 1.4.4.7. Place the PCR tube on a magnetic rack for 2 minutes to allow the solution to clarify.

[0266] 1.4.4.8. Use a pipette to draw 20 μl of supernatant and transfer it to a new PCR tube.

[0267] 1.5 Treatment and purification of bisulfite:

[0268] 1.5.1. Prepare the required reagents in advance and dissolve them. Add the reagents according to the table below:

[0269]

[0270] 1.5.2. DNA Protect buffer turns the liquid blue upon addition. Gently pipette to mix, then divide into two tubes and place them on the PCR instrument.

[0271] 1.5.3. Configure and run the following program:

[0272] Heat cover 105℃.

[0273]

[0274] 1.5.4. Brief centrifugation: Combine the two identical samples into a single clean 1.5 ml centrifuge tube.

[0275] 1.5.5. Add 310 μl of Buffer BL to each sample (add 1 μl of Carrier RNA (1 μg / μl) for sample volumes less than 100 ng), vortex to mix, and briefly centrifuge.

[0276] 1.5.6. Add 250 μl of anhydrous ethanol to each sample, vortex to mix for 15 s, centrifuge briefly, and add the mixture to the corresponding prepared centrifuge column.

[0277] 1.5.7. Let stand for 1 minute, centrifuge for 1 minute, transfer the liquid in the collection tube back to the centrifuge column, centrifuge for 1 minute, and discard the liquid in the centrifuge tube.

[0278] 1.5.8. Add 500 μl of buffer BW (note whether to add anhydrous ethanol), centrifuge for 1 min, and discard the waste liquid.

[0279] 1.5.9. Add 500 μl of buffer BD (note whether to add anhydrous ethanol), cap the tube, and incubate at room temperature for 15 min. Centrifuge for 1 min and discard the liquid collected after centrifugation.

[0280] 1.5.10. Add 500 μl buffer BW (note whether to add anhydrous ethanol), centrifuge for 1 min, discard the separated liquid, and repeat once, for a total of 2 times.

[0281] 1.5.11. Add 250 μl of anhydrous ethanol, centrifuge for 1 min, transfer the centrifuge column to a new 2 ml collection tube, and discard all remaining liquid.

[0282] 1.5.12. Place the centrifuge column into a clean 1.5ml centrifuge tube, add 20μl of nuclease-free water to the center of the centrifuge column membrane, gently cap the tube, incubate at room temperature for 1 min, and centrifuge for 1 min.

[0283] 1.5.13. Transfer the liquid in the collection tube back to the centrifuge column, let it stand at room temperature for 1 min, and then centrifuge for 1 min.

[0284] 1.6 Pre-amplification and purification before hybridization:

[0285] 1.6.1. Prepare the reaction system according to the table below, mix by pipetting, and briefly centrifuge:

[0286]

[0287] 1.6.2. Set up the following program and start the PCR program:

[0288] Heat cap 105℃

[0289]

[0290] 1.6.3. The number of PCR cycles should be adjusted according to the amount of DNA used. Reference data is shown below:

[0291]

[0292] 1.6.4. Add 50 μl of AgencourtAMPure XP magnetic beads to the PCR tube after the reaction is complete, and mix well by pipetting to avoid generating air bubbles (AgencourtAMPure XP should be mixed and equilibrated at room temperature beforehand).

[0293] 1.6.5. Incubate at room temperature for 5-15 minutes, then place the PCR tube on a magnetic rack for 3 minutes to allow the solution to clarify.

[0294] 1.6.6. Remove the supernatant, keep the PCR tube on the magnetic rack, add 200 μl of 80% ethanol solution to the PCR tube, and let it stand for 30 seconds.

[0295] 1.6.7. Remove the supernatant, then add 200 μl of 80% ethanol solution to the PCR tube, let it stand for 30 seconds, and then completely remove the supernatant (it is recommended to use a 10 μl pipette to remove any residual ethanol solution at the bottom).

[0296] 1.6.8. Let stand at room temperature for 5 minutes to allow the residual ethanol to evaporate completely.

[0297] 1.6.9. Add 30 μl of Nuclease-free water, remove the centrifuge tube from the magnetic rack, and gently aspirate and resuspend the magnetic beads using a pipette.

[0298] 1.6.10. Let stand at room temperature for 2 min, then place the 200 μl PCR tube on a magnetic rack for 2 min to allow the solution to clarify.

[0299] 1.6.11. Use a pipette to transfer the supernatant to a new 200 μl PCR tube (place on an ice box), label the sample number on the reaction tube, and prepare for the next reaction.

[0300] 1.6.12. Take 1 μl of sample and use Qubit to determine the library concentration, and record the library concentration.

[0301] 1.6.13. Take 1 μl of sample and use Agilent 2100 to determine the length of the library fragments. The library length is approximately between 270 bp and 320 bp.

[0302] 1.6.14. Sequencing was performed using the Illumina high-throughput sequencing platform.

[0303] 1.6.15. Bioinformatics analysis workflow for methylation and tumor data.

[0304] The process is as follows: Quality control software such as FASTP is used to check the quality of the raw sequencing data, and low-quality reads are filtered, truncated, or removed to obtain the corresponding clean data; the quality-controlled clean data is aligned to the reference genome (hg19) using Bismark Bowtie2 alignment software; duplicates in the initial alignment BAM file are removed using deduplicate_bismark; the corresponding methylation site information is extracted using Bismark_methylation_extractor to obtain the final methylated CG file (including all individual CG site information files); the sliding window calculates the methylation level of the target region, ichor-CNA is used to calculate the tumor fraction and tumor ploidy, cfDNA fragment size is used to predict the tumor probability, and read distribution is used to calculate the tumor proportion.

[0305] Methylation level: A sliding window is used. For each window, the number of CG sites is counted. Since the methylation depth of cytosine at each CG site and the total depth of the site are known, the methylation level of the entire window can be calculated. This is the ratio of the sum of the methylation depths of cytosine at all CG sites to the sum of the total depths of all CG sites. Each window will have a corresponding methylation level calculated using the above method. The methylation depth of cytosine at each CG site is the number of reads that show methylated cytosine at that site in the sequencing results, i.e., the number of reads that show C (cytosine) at that site. The total depth of the site is the total number of all sequencing reads covering that site, i.e., the total number of reads that show C or T (thymine) at that site.

[0306] Tumor fraction and tumor ploidy: The ichorCNA software was used to process the BAM file. A Hidden Markov Model (HMM) was used to calculate the CNA, and the Expectation-Maximization (EM) algorithm was used to calculate the tumor fraction and tumor ploidy. The detailed working principle of the ichorCNA software is as follows:

[0307] (1) First, set the bin size (10M or 5M) and calculate the number of reads for each bin;

[0308] (2) Correct the number of reads for each bin for GC, mappability, and depth differences;

[0309] (3) Calculate the logR value by comparing the number of reads in each bin after correction with the number of reads in each bin of the software's built-in normal panel;

[0310] (4) Calculate the result of each possible scheme and the maximum likelihood function of each scheme using the HMM and EM algorithms, where the HMM and EM algorithms are the internal models of the iChorCNA software;

[0311] (5) Select the scheme with the largest likelihood function as the final result, that is, obtain the tumor fraction and tumor ploidy.

[0312] cfDNA fragment length calculation: For the obtained BAM file, reads with MAPQ < 30 are filtered out, and whole-genome fragment information is extracted using the R package GCcontent. The extracted fragment information is then laid out into adjacent, non-overlapping 100kb intervals based on the autosomes of the hg19 reference genome. Short fragments are defined as short (150-200bp), long fragments as long (201-320bp), and cov fragments are short + long. The number of short, long, and cov fragments in each interval is counted, representing the fragment coverage. The LOWESS algorithm is used to correct for short, long, and cov fragments, removing coverage bias caused by GC offsets. The 100kb intervals are then merged into 5MB intervals, resulting in 499 non-overlapping intervals. The number of short, long, and cov fragments in each 5MB interval is counted. The short / long ratio represents the short-to-long fragment ratio. Through analysis, cov (499 intervals corresponding to 499 cov features) was selected as the feature, PCA was used for dimensionality reduction, and then a gradient boosting tree model was used to obtain the cancer risk probability value.

[0313] The cfDNA fragment length feature value analysis module calculates the risk feature value of tumor related to the length of the cfDNA fragment extracted from the sequencing data obtained from the sequencing module. This is based on the ratio between short and long fragments in the cfDNA fragment length, and the cfDNA fragment length feature value is calculated through a gradient boosting tree model.

[0314] Preferably, the long segment is a segment with a length of 201~320bp;

[0315] Short segments are segments with a length of 150~200bp;

[0316] Reads distribution: Using the BAM file, the CancerDecteor software is used to calculate the likelihood value of a given sequencing fragment in the test sample originating from one of the three source components: healthy plasma, normal tissue, and tumor tissue. This yields the proportion of reads in the test sample originating from the tumor-derived component. In other words, it represents the probability that a given read originates from one of the three source components: healthy plasma, normal tissue, or tumor tissue.

[0317] Example 2

[0318] This embodiment classifies 58 samples (37 healthy individuals and 21 lung cancer patients), dividing them into a 70% training set and a 30% test set. The training set consists of 15 lung cancer patient cfDNA samples and 26 healthy individuals, while the test set consists of 6 lung cancer patients and 11 healthy individuals. See [link to relevant documentation]. Figure 1 Based on the training set described above, and according to the method described in Example 1, multiple lung cancer-related biomarkers (i.e., methylation level, cfDNA fragment length feature value, tumor score, tumor ploidy, and read distribution) are integrated, and the GLM algorithm is used to obtain the MODE model. The specific calculation process is shown below.

[0319] Table 1: Numerical values ​​of methylation level, cfDNA fragment length features, tumor score, tumor ploidy, and read distribution in the training set.

[0320]

[0321] The values ​​of the five markers in Table 1 were imported into R software and integrated using the GLM algorithm. The response variable follows a binomial distribution, the link function is logistic regression, and the regression coefficients were solved by 100 iterations according to the maximum likelihood estimation principle. Finally, the MODE model was fitted, as shown in Equation 1.

[0322]

[0323] (Formula 1)

[0324] Model Interpretation: P The probability of an outcome when exposed to a certain state. logit ( P () is a variable transformation method, representing the transformation of... P conduct logit Transformation, β These are partial regression coefficients, representing the coefficients for each unit change in independent variables, assuming other independent variables remain constant. logit ( P The estimated value of ).

[0325] P represents the probability that the subject will develop cancer;

[0326] X1 represents the tumor score;

[0327] X2 represents the ploidy of the tumor;

[0328] X3 represents the methylation level;

[0329] X4 is the cfDNA fragment length feature value;

[0330] X5 represents the reads distribution;

[0331] β0 is any value selected from 100 to 150, preferably 114.12;

[0332] β1 is any value selected from -700 to 0, preferably -686.25;

[0333] β2 is any value selected from 20 to 40, preferably 31.16;

[0334] β4 is any value selected from -20 to 0, preferably -17.45;

[0335] Among them, the matrix analysis based on the methylation level of the target region yields Equation 2 from Table 1:

[0336] β3X3= a* X 31 + b* X 32+ c* X 33 + d* X 34 + e* X 35 + f* X 36 + g* X 37 (Formula 2)

[0337] X 31 The methylation level at positions 223721500-223726500 on chromosome 2 in the target region;

[0338] X 32 The methylation level at positions 170147000-170152000 on chromosome 6 in the target region;

[0339] X 33 The methylation level of chromosome 8 from position 182500 to 187500 in the target region;

[0340] X 34 The methylation level at positions 64081500-64086500 on chromosome 8 in the target region;

[0341] X 35 The methylation level at positions 14688500-14693500 on chromosome 9 in the target region;

[0342] X 36 The methylation level of chromosome 10 from positions 119803500 to 119808500 in the target region;

[0343] X 37 The methylation level at positions 56285500-56290500 on chromosome 15 in the target region;

[0344] a can be any value selected from -300 to 300, preferably -260.66;

[0345] b is any value selected from -200 to 200, preferably -181.95;

[0346] c is any value selected from -100 to 100, preferably -85.21;

[0347] d is any value selected from -350 to 350, preferably -305.36;

[0348] e can be any value selected from -250 to 250, preferably 218.80;

[0349] f is any value selected from -100 to 100, preferably 60.85;

[0350] g is any value selected from -250 to 250, preferably -209.25.

[0351] Among them, the matrix analysis based on the reads distribution yields Equation 3 from Table 1:

[0352] β5X5=l* X 51 + m* X 52+ n* X 53 (Formula 3)

[0353] X 51 Distribution of reads derived from healthy plasma;

[0354] X 52 The distribution of reads originating from normal tissues;

[0355] X 53 Distribution of reads derived from tumor tissue;

[0356] l can be any value selected from 50 to 200, preferably 166.02;

[0357] m can be any value selected from 0 to 10, preferably 8.41;

[0358] n is any value selected from -0.1 to 1, preferably 0.00000001.

[0359] in,

[0360]

[0361] The `prediction` function in R can be used to obtain the probability values ​​predicted by the MODE model, thus determining that the AUC value of the model on the test set (6 lung cancer patients and 11 healthy individuals) is 1. Figure 1 As shown.

[0362] Example 3

[0363] Based on the results of Example 1, following the method of Example 2, a generalized linear regression model was constructed using tumor fraction, tumor ploidy, read distribution, cfDNA fragment size, and a single biomarker of methylation level on the training set (cfDNA from 15 lung cancer patients and 26 healthy individuals). The AUCs obtained on the test set (6 lung cancer patients and 11 healthy individuals) were 0.788, 0.636, 0.712, 0.955, and 0.985, respectively. The results are as follows: Figures 2-6 As shown, the specific modeling process is the same as in Example 2.

[0364] Example 4

[0365] Based on the results of Example 1, following the method of Example 2, a generalized linear regression model was constructed using two biomarkers, read distribution and tumor ploidy, on the training set (cfDNA from 15 lung cancer patients and 26 healthy individuals). The AUC obtained on the test set (6 lung cancer patients and 11 healthy individuals) was 0.788. The results are as follows... Figure 7 As shown, the specific modeling process is the same as in Example 2.

[0366] Example 5

[0367] Based on the results of Example 1, following the method of Example 2, a model was constructed using a generalized linear regression model on the training set (cfDNA from 15 lung cancer patients and 26 healthy individuals) using a combination of three biomarkers: tumor ploidy, tumor fraction, and cfDNA fragment size. The AUC obtained on the test set (6 lung cancer patients and 11 healthy individuals) was 0.970. The results are as follows... Figure 8 As shown, the specific modeling process is the same as in Example 2.

[0368] In Example 2, the MODE model obtained by simultaneously using five biomarkers (tumor fraction, tumor ploidy, reads distribution, cfDNA fragment size feature value, and methylation level) achieved an AUC of 1 on the test set, which is higher than the AUC obtained by modeling using the same method in Examples 3, 4, and 5. Furthermore, the AUC obtained by combining multiple biomarkers is also higher than the AUC obtained by modeling with a single biomarker, indicating that using multiple biomarkers improves classification performance in this model. These results also demonstrate that by integrating five biomarkers (tumor fraction, tumor ploidy, reads distribution, cfDNA fragment size feature value, and methylation level) to obtain the MODE model, the best prediction and classification performance can be achieved based on information covering multiple dimensions such as methylation, fragmentation, and tumor fraction.

Claims

1. A system for lung cancer screening, comprising: The data acquisition module is used to acquire data on methylation levels, tumor scores, tumor ploidy, cfDNA fragment length characteristics, and read distribution in the target region of the subject; and The lung cancer calculation module predicts the probability of a subject developing lung cancer based on data obtained from the data acquisition module, including methylation level, tumor score, tumor ploidy, cfDNA fragment length characteristics, and read distribution. The methylation level of the target region is calculated based on the methylation level of each CG site in the target region, wherein the methylation level of the CG site is the ratio of the methylated cytosine detected at that site to the sum of the unmethylated cytosine and the unmethylated cytosine in all detected sequence results of that site; Tumor fraction refers to the proportion of ctDNA in cfDNA, that is, the proportion of cell-free DNA released by tumor cells in the total cfDNA; Tumor ploidy refers to the variation caused by abnormalities in chromosome structure and number, that is, the number of ploids in a tumor; The cfDNA fragment length feature value is based on the sequencing data obtained from the sequencing module. The length information of the cfDNA fragment is extracted, and the risk feature value of lung cancer related to the cfDNA fragment length is calculated by using a gradient boosting tree model based on the ratio between short fragments and long fragments in the cfDNA fragment length. Long segments are segments with a length of 201~320bp; Short segments are segments with a length of 150~200bp; The reads distribution refers to the probability of the sample origin obtained by statistically analyzing each read in the cfDNA sequencing data, that is, the probability that a given read comes from one of three source components: healthy plasma, normal tissue, and tumor tissue. The target area includes the following areas: Chromosome 2, positions 223721500-223726500. Chromosome 6, positions 170147000-170152000. Chromosome 8, positions 182500-187500, Chromosome 8, positions 64081500-64086500. Chromosome 9, positions 14688500-14693500, Chromosome 10, positions 119803500-119808500, and Chromosome 15, positions 56285500-56290500.

2. The system according to claim 1, wherein, The data acquisition module includes a sequencing module, a methylation level analysis module, a tumor score analysis module, a tumor ploidy analysis module, a cfDNA fragment length feature value analysis module, and a reads distribution analysis module. The sequencing module is used to perform whole-genome sequencing on the subject's cfDNA; The methylation level analysis module analyzes the methylation level of the target region based on sequencing data obtained from the sequencing module. The tumor score analysis module analyzes tumor scores based on sequencing data obtained from the sequencing module; The tumor ploidy analysis module analyzes tumor ploidy based on sequencing data obtained from the sequencing module; The cfDNA fragment length feature value analysis module calculates a risk feature value for tumors related to the length of the cfDNA fragment extracted from the sequencing data obtained from the sequencing module. The reads distribution analysis module analyzes the reads distribution from the sequencing data obtained from the sequencing module and then predicts the source of each read.

3. The system according to claim 1, wherein, cfDNA sequencing data is cfDNA sequencing data after removing low-quality sequencing fragments.

4. The system according to claim 3, wherein, cfDNA sequencing data are sequencing data that have been further excluded after removing low-quality sequencing fragments and eliminating sequencing data in the low alignment rate range.

5. The system according to claim 1, wherein, The lung cancer calculation module pre-stores a formula fitted based on known sample methylation levels, tumor fractions, tumor ploidy, cfDNA fragment length characteristics, and read distribution to predict the probability of a subject developing lung cancer. This formula is Equation 1. Set 1 Mode P represents the probability that the subject will develop lung cancer; X1 represents the tumor score; X2 represents the ploidy of the tumor; X3 represents the methylation level; X4 is the cfDNA fragment length feature value; X5 represents the reads distribution; β0 is any value selected from 100 to 150; β1 is any value selected from -700 to 0; β2 is any value selected from 20 to 40; β4 is any value selected from -20 to 0.

6. The system according to claim 5, wherein, β0 is 114.12; β1 is -686.25; β2 is 31.16; β4 is -17.

45.

7. The system according to claim 5, wherein, β3X3 is obtained from Equation 2: β3X3 = a*X31 + b*X32 + c*X33 + d*X34 + e*X35 + f*X36 + g*X37 (Equation 2) X31 represents the methylation level at positions 223721500-223726500 on chromosome 2 within the target region; X32 represents the methylation level at positions 170147000-170152000 on chromosome 6 within the target region; X33 represents the methylation level at positions 182500-187500 on chromosome 8 within the target region; X34 represents the methylation level at positions 64081500-64086500 on chromosome 8 within the target region; X35 represents the methylation level at positions 14688500-14693500 on chromosome 9 within the target region; X36 represents the methylation level at positions 119803500-119808500 on chromosome 10 within the target region; X37 represents the methylation level at positions 56285500-56290500 on chromosome 15 within the target region; a is any value selected from -300 to 300; b is any value selected from -200 to 200; c is any value selected from -100 to 100; d is any value selected from -350 to 350; e is any value selected from -250 to 250; f is any value selected from -100 to 100; g is any value selected from -250 to 250.

8. The system according to claim 7, wherein, a is -260.66; b is -181.95; c is -85.21; d is -305.36; e is 218.80; f is 60.85; g is -209.

25.

9. The system according to claim 5, wherein, β5X5 is obtained from Equation 3: β5X5=l* X 51 + m* X 52+ n* X 53 Formula 3 X 51 Distribution of reads derived from healthy plasma; X 52 The distribution of reads originating from normal tissues; X 53 Distribution of reads derived from tumor tissue; l is any value selected from 50 to 200; m is any value selected from 0 to 10; n is any value selected from -0.1 to 1.

10. The system according to claim 9, wherein, l is 166.02; m is 8.41; n is 0.00000001.

11. The system according to claim 1, wherein, The system also includes a bisulfite treatment module for treating the subject's cfDNA with bisulfite.

12. A method for lung cancer screening using any one of the systems of claims 1-11, the method comprising: The sample collection process involves acquiring data on methylation levels, tumor scores, tumor ploidy, cfDNA fragment length characteristics, and read distribution in the target region of the subject. The lung cancer calculation step uses data on methylation level, tumor score, tumor ploidy, cfDNA fragment length, and read distribution obtained from the data acquisition module to predict the probability of a subject developing lung cancer.

13. The method for lung cancer screening according to claim 12, the method comprising: The sample collection steps include sequencing, methylation level analysis, tumor fraction analysis, tumor ploidy analysis, cfDNA fragment length characteristic value analysis, and reads distribution analysis. The sequencing step is used to perform whole-genome sequencing of the subject's cfDNA; The methylation level analysis step is based on sequencing data obtained from the sequencing module to analyze the methylation level of the target region. The tumor score analysis step is based on analyzing tumor scores using sequencing data obtained from the sequencing module; The tumor ploidy analysis step is based on analyzing tumor ploidy using sequencing data obtained from the sequencing module; The cfDNA fragment length feature value analysis step calculates a length-related risk feature value for tumors based on the length of the cfDNA fragment extracted from the sequencing data obtained from the sequencing module. The reads distribution analysis step is based on analyzing the reads distribution from the sequencing data obtained from the sequencing step, thereby predicting the source of each read.

Citation Information

Patent Citations

  • Peripheral blood methylation gene and IDH1 combined detection and diagnosis model for lung cancer

    CN111172279A

  • Kit for early screening of hepatocellular carcinoma and preparation method and application thereof

    CN111690740A

  • Methods for cancer detection and monitoring

    US20190316184A1