Methods for predicting cancer risk based on cell-free DNA and artificial intelligence
Capillary electropherogram analysis of cfDNA features with AI predicts cancer probability, addressing the limitations of current methods by enhancing sensitivity, accuracy, and reducing costs for early detection and monitoring.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SEEKIN INC SHENZHEN CHINA
- Filing Date
- 2025-04-27
- Publication Date
- 2026-04-30
AI Technical Summary
Current cancer screening methods are costly, complex, and limited to specific cancer types, making them unsuitable for widespread use, especially in low-resource settings, and there is a need for more effective and affordable methods for early detection, recurrence monitoring, and treatment response assessment.
A method using capillary electropherogram to analyze cell-free DNA (cfDNA) features such as concentration and short fragment ratio from a mono-nucleosome, combined with protein tumor markers and artificial intelligence, to predict cancer probability with improved sensitivity, accuracy, and cost-effectiveness.
Enables sensitive, accurate, and cost-effective cancer detection and monitoring by leveraging multidimensional cfDNA and protein markers, reducing complexity and infrastructure requirements.
Smart Images

Figure CN2025091413_30042026_PF_FP_ABST
Abstract
Description
METHODS FOR PREDICTING CANCER RISK BASED ON CELL-FREE DNA AND ARTIFICIAL INTELLIGENCETECHNICAL FIELD
[0001] The present disclosure relates to methods of determining the concentration of cell-free DNA (cfDNA) and the short fragment ratio from a mono-nucleosome by capillary electropherogram. Specifically, the present application relates to a method, system, electronic device, and computer-readable medium for predicting a probability that a sample to be tested is derived from a cancer patient based on plasma features and artificial intelligence.BACKGROUND
[0002] Many studies have indicated that circulating tumor DNA (ctDNA) fragments from tumor cells in the blood are shorter than normal cell-free DNA (cfDNA) , and the size of cfDNA fragments can be assessed by Next-Generation Sequencing (NGS) approach. Meanwhile, the fragmentation pattern of cfDNA in the genome is significantly different between healthy subjects and cancer patients, and also different between different cancer types. The current standard-of-care (SOC) cancer screening modalities including imaging, plasma tumor markers, and cytology are restricted to particular cancer types and have unsatisfactory accuracy and participant compliance. In addition, current methods heavily rely on NGS to detect fragmentation patterns of cfDNA, which is complex and costly. Thus, there is a need to develop more effective and affordable methods that can be utilized for cancer early detection, recurrence monitoring, treatment response assessment as well as mechanistic study of the cause of individual cancers.SUMMARY
[0003] The present disclosure solves one of the technical problems in the related field. In this regard, the present disclosure provides a non-invasive method for cancer detection, recurrence monitoring, and treatment response assessment based on multidimensional characteristics of cell-free DNA (cfDNA) and / or protein markers in plasma and artificial intelligence, based on a technical route of cancer genome panorama in combination with protein tumor markers. This technology is based on a capillary electropherogram to decode the plasma features. At the same time, in combination with specific protein tumor markers as well as big data and artificial intelligence, it can predict a probability that the sample to be tested is derived from a cancer patient. Based on multiple features (including the concentration of cfDNA and the short fragment ratio from a mono-nucleosome) of the sample to be tested, the present disclosure employs a multidimensional and multivariable weighting algorithm and combines cfDNA features (e.g., based on concentration and short fragment ratio of cfDNA from a mono-nucleosome) and / or protein tumor markers (e.g., based on levels of a panel of protein biomarkers in blood) , such that the probability that the sample to be tested is derived from a cancer patient can be predicted in a more sensitive, more accurate, and specific manner under the premise of more controllable testing costs. Compared with NGS-based technologies, this detection methods described herein are less complex and more cost-effective.
[0004] In one aspect, the disclosure is related to a method for determining the concentration of cell-free DNA (cfDNA) within a size range in a test sample, the method comprising: (a) isolating cfDNA from the test sample; (b) analyzing the cfDNA by capillary electropherogram to generate a distribution curve, in some embodiments, a ladder and / or a control sample are analyzed in the same batch of the cfDNA, in some embodiments, the control sample comprises a pre-determined DNA concentration, in some embodiments, the ladder comprises one or more markers, and at least one marker from the one or more markers has a pre-determined DNA length; and (c) determining the concentration of the cfDNA within the size range using the control sample and / or the ladder as references. In some embodiments, the control sample is a circulating tumor DNA reference sample prepared by inducing apoptosis in tumor cells. In some embodiments, the distribution curve exhibits an upward or downward shift. In some embodiments, the ladder comprises a first marker and a second marker, and the size range is pre-defined by the first marker and the second marker. In some embodiments, the size range is determined by analyzing the distribution curve of the cfDNA. In some embodiments, the size range is defined by one or more local maximum points and / or one or more local minimum points of the distribution curve (e.g., one local maximum point and a local minimum point of the distribution curve) . In some embodiments, the size range is defined by two neighboring local minimum points of the distribution curve. In some embodiments, the size range is defined by one local maximum point or one local minimum point, and a pre-defined ladder marker. In some embodiments, the concentration of cfDNA is determined by measuring the size of an area defined by the distribution curve within the size range and a straight line connecting the two neighboring local minimum points of the distribution curve, and optionally the size of the area is calibrated by measuring the area size within a reference size range (e.g., the size range) from the control sample and / or the ladder across different batches, in some embodiments, the control sample and / or the ladder have the same DNA concentration among the different batches. In some embodiments, the size range corresponds to a mono-nucleosome.
[0005] In one aspect, the disclosure is related to a method for determining the short fragment ratio of cfDNA from a mono-nucleosome in a test sample, the method comprising: (a) isolating cfDNA from the test sample; (b) analyzing the cfDNA by capillary electropherogram to generate a distribution curve, in some embodiments, a control sample and / or a ladder are analyzed in the same batch of the cfDNA, in some embodiments, the control sample comprises a pre-determined DNA concentration, in some embodiments, the ladder comprises one or more markers, and at least one marker from the one or more markers has a pre-determined DNA length; and (c) determining the short fragment ratio of cfDNA from the mono-nucleosome using the control sample and / or the ladder as references. In some embodiments, the control sample is a circulating tumor DNA reference sample prepared by inducing apoptosis in tumor cells. In some embodiments, the distribution curve exhibits an upward or downward shift. In some embodiments, step (c) comprises identifying a size range of the cfDNA that corresponds to the mono-nucleosome; determining a first area size defined by the distribution curve and a straight line connecting two neighboring local minimum points of the distribution curve of the cfDNA within the size range; determining a second area size defined by the distribution curve, a pre-defined ladder marker representing the upper size range of the short fragment (e.g., 150 bp) , and the straight line connecting the two neighboring local minimum points of the distribution curve of the cfDNA within the size range, and optionally the first and second area sizes are calibrated by measuring the area size within a reference size range (e.g., the size range of the cfDNA that corresponds to the mono-nucleosome) from the control sample and / or the ladder across different batches, in some embodiments, the control sample and / or the ladder have the same DNA concentration among the different batches; and calculating the short fragment ratio of cfDNA from the mono-nucleosome by dividing the second area size by the first area size. In some embodiments, the method further comprises determining a probability that the test sample is derived from a cancer patient based on the concentration of the cfDNA within a size range that corresponds to a mono-nucleosome and / or the short fragment ratio of cfDNA from the mono-nucleosome.
[0006] In one aspect, the disclosure is related to a computer-implemented method for early detection of the presence of cancer in a subject, the method comprising: (a) determining the concentration of cfDNA within a size range corresponding to a mono-nucleosome in a test sample of the subject; (b) determining the short fragment ratio of cfDNA from the mono-nucleosome in the test sample from the subject; (c) selecting a plurality of parameters for inputs into a machine learning system, in some embodiments, the plurality of parameters comprises the concentration of cfDNA and the short fragment ratio of cfDNA from the mono-nucleosome; (d) training the machine learning system using a machine learning algorithm selected from Random Forest (RF) , Generalized Linear Model (GLM) , Support Vector Machine (SVM) , and / or Gradient Boosting Machine (GBM) ; and (e) determining a cancer predicting score, in some embodiments, a high cancer predicting score indicates a high probability of the subject to have cancer.
[0007] In one aspect, the disclosure is related to a computer-implemented method for early detection of the presence of cancer in a subject, the method comprising: (a) quantifying the level of a panel of biomarkers from a test sample (e.g., a blood sample) of the subject, in some embodiments, the panel of biomarkers comprises one or more protein biomarkers selected from AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA21-1, SCCA, ProGRP, and PSA; (b) determining the concentration of cfDNA within a size range corresponding to a mono-nucleosome in the test sample of the subject; (c) determining the short fragment ratio of cfDNA from the mono-nucleosome in the test sample of the subject; (d) selecting a plurality of parameters for inputs into a machine learning algorithm system, in some embodiments, the plurality of parameters comprises the level of the one or more protein biomarkers, the concentration of cfDNA within the size range corresponding to the mono-nucleosome, and the short fragment ratio of cfDNA from the mono-nucleosome; (e) training the machine learning system using a machine learning algorithm selected from Random Forest (RF) , Generalized Linear Model (GLM) , Support Vector Machine (SVM) , and / or Gradient Boosting Machine (GBM) ; and (f) determining a cancer predicting score, in some embodiments, a high cancer predicting score indicates a high probability of the subject to have cancer. In some embodiments, the plurality of parameters further comprises at least one clinical parameter (e.g., age, gender, and / or smoking status) . In some embodiments, the concentration of cfDNA within the size range corresponding to the mono-nucleosome and the short fragment ratio of cfDNA from the mono-nucleosome are determined by capillary electropherogram. In some embodiments, the concentration of cfDNA within the size range corresponding to the mono-nucleosome and the short fragment ratio of cfDNA from the mono-nucleosome are determined by isolating cfDNA from the test sample; analyzing the cfDNA by capillary electropherogram to generate a distribution curve, in some embodiments, a control sample and / or a ladder are analyzed in the same batch of the cfDNA; and determining the concentration of cfDNA within the size range corresponding to the mono-nucleosome and the short fragment ratio of cfDNA from the mono-nucleosome using the control sample and / or the ladder as references. In some embodiments, the method can aid early detection of the presence of at least two cancer types simultaneously.
[0008] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Methods and materials are described herein for use in the present invention; other, suitable methods and materials known in the art can also be used. The materials, methods, and examples are illustrative only and not intended to be limiting. All publications, patent applications, patents, sequences, database entries, and other references mentioned herein are incorporated by reference in their entirety. In case of conflict, the present specification, including definitions, will control.
[0009] Other features and advantages of the invention will be apparent from the following detailed description and figures, and from the claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0010] FIG. 1A shows the cfDNA fragment size distribution based on Next-Generation Sequencing (NGS) .
[0011] FIG. 1B shows the cfDNA fragment size distribution based on capillary electropherogram. The result was obtained using an Agilent 2100 Bioanalyzer. The starting position ( “Start” ) , the 150 bp position ( “150bp” ) , and the ending position ( “End” ) of a mono-nucleosome are labeled. The Y-axis represents the fluorescence intensity of the labeled nucleic acids, which correlates with their concentrations in the sample. “FU” stands for fluorescence unit. The X-axis represents the migration time.
[0012] FIG. 2 shows the overlaid profiles of the result of one sample's electropherogram (black line) and the ladder result (grey line) . The results were obtained using Agilent 2100 Bioanalyzer software (Agilent Technologies Inc) . The 150 bp peak ( “150bp” ) , the 200 bp peak ( “200bp” ) , and the 300 bp peak ( “300bp” ) of the ladder are labeled. The lowest point between the 150 bp peak and the 300 bp peak represents the ending position of the mono-nucleosome. The Y-axis represents peak fluorescence intensity and the X-axis represents the migration time.
[0013] FIG. 3 shows the overlaid profiles of different batches of ladders. The Y-axis represents the peak fluorescence intensity and the X-axis represents the migration time. A zoomed-in result of peaks migrated before 90 seconds is shown on the right side. The horizontal and vertical shiftings demonstrated the necessity of inter-batch correction.
[0014] FIG. 4 shows one sample's electropherogram profile with a downward shift. A fitted straight line (shown as a dashed line) can be used as the baseline for calculating the area under the curve.
[0015] FIG. 5 shows boxplot results of the optimized cfDNA concentrations from a mono-nucleosome in cancer patients ( “Cancer” ) and healthy individuals ( “Healthy” ) .
[0016] FIG. 6 shows the ROC (receiver-operating characteristic) curves indicating the performance of cancer detection using different methods of determining cfDNA concentration. The optimized cfDNA concentrations from a mono-nucleosome (cfDNA v2) shows an improved performance than the conventional methods (cfDNA v0) .
[0017] FIG. 7A shows boxplot results of the proportion of cfDNA short fragments (those smaller than 150 bp) in cancer patients ( “Cancer” ) and healthy individuals ( “Healthy” ) .
[0018] FIG. 7B shows a ROC curve indicating the performance of cancer detection based on the proportion of cfDNA short fragments (those smaller than 150 bp) that was determined using the optimized cfDNA concentrations from a mono-nucleosome (P150 v2) .
[0019] FIG. 8 shows the complementary relationship between the cfDNA concentration from a mono-nucleosome (cfDNA v2) and the short fragment ratio (P150 v2) in predicting true positive samples. The numbers in the Venn diagram represent the count of true positive cases detected by cfDNA v2, P150 v2, or both.
[0020] FIG. 9A shows boxplot results of the probability of predicting cancer obtained using a machine learning model that combined the cfDNA concentration from a mono-nucleosome (cfDNA v2) and the short fragment ratio (P150 v2) in cancer patients ( “Cancer” ) and healthy individuals ( “Healthy” ) .
[0021] FIG. 9B shows a ROC curve indicating the performance for distinguishing cancer patients from non-cancer individuals using the probability of predicting cancer obtained using the machine learning model in FIG. 9A.
[0022] FIG. 10A shows boxplot results of the probability of predicting cancer obtained using a machine learning model that combined the cfDNA concentration from a mono-nucleosome (cfDNA v2) , the short fragment ratio (P150 v2) , and seven protein biomarkers (AFP, CA125, CA153, CA199, CA724, CEA, CYFRA21-1) in cancer patients ( “Cancer” ) and healthy individuals ( “Healthy” ) .
[0023] FIG. 10B shows a ROC curve indicating the performance for distinguishing cancer patients from non-cancer individuals using the probability of predicting cancer obtained using the machine learning model in FIG. 10A.DETAILED DESCRIPTION
[0024] Cancer is an important public health issue worldwide. The global cancer burden is increasing rapidly, and nearly 19.3 million new cases and 10.0 million cancer deaths were estimated in 2020. It is estimated that more than two-thirds of annual cancer deaths in the world occur in low-or middle-income countries (LMICs) . The global cancer burden is expected to be 28.4 million cases in 2040, a47%rise from 2020, with a larger increase in transitioning (64%to 95%) versus transitioned (32%to 56%) countries due to demographic changes, although this may be further exacerbated by increasing risk factors associated with limited medical infrastructure in LMICs. It is well acknowledged that cancer early detection offers a higher cure rate and 5-year survival rate as well as a reduction in treatment cost and loss of economic productivity. Various cancer screening techniques are currently available in clinical practice. Examples include low-dose computed tomography (LDCT) for lung cancer screening, mammogram used to detect breast cancer, HPV test or cytology combined with colposcopy for early detection of cervical cancer, faecal occult blood test (FOBT) combined with colonoscopy for colorectal cancer screening, and prostate-specific antigen (PSA) for prostate cancer. However, the high cost of these screening methods and their need for specialised infrastructure and skilled technicists limit their application, and this is why there are alternatives of screening tests in LMICs: visual inspection with acetic acid (VIA) for cervical cancer and clinical breast examination (CBE) for breast cancer screening. Moreover, these methods are individually designed for screening for specific cancer types, hindering their widespread use as screening tools. In addition, liquid biopsy methods, which detect blood-based analytes such as cancer-derived DNA, are now being adopted in human medicine to simultaneously screen for multiple types of cancer; these MCED tests represent a paradigm shift in cancer screening and promise to significantly increase the number of patients with cancer that are detected at earlier stages. However, these tests are not suitable for using in LMICs due to cost, complexity, and dependency on high-end infrastructure and a rigorous laboratory. Taken together, these factors contribute to the fact that cancer is often diagnosed at a later stage and to inequalities in health care. In order to perform large-scale cancer screening among apparently healthy individuals in the future, especially among the population in LMICs, the development and validation of a more general, robust, and affordable MCED test are essential.
[0025] The present disclosure discloses methods of determining the concentration of cell-free DNA (cfDNA) and methods of determining the short fragment ratio from a mono-nucleosome by capillary electropherogram. These two cfDNA features, optionally in combination with the levels of a selected panel of protein biomarkers in blood, can be used as inputs to train a machine learning system thereby predicting cancer in a more sensitive, more accurate, and more cost-effective way.
[0026] Capillary Electropherogram and its Application on DNA Fragment Analysis
[0027] Capillary electrophoresis (CE) is a family of electrokinetic separation methods performed in submillimeter diameter capillaries and in micro-and nanofluidic channels. Very often, CE refers to capillary zone electrophoresis (CZE) , but other electrophoretic techniques including capillary gel electrophoresis (CGE) , capillary isoelectric focusing (CIEF) , capillary isotachophoresis and micellar electrokinetic chromatography (MEKC) belong also to this class of methods. In CE methods, analytes migrate through electrolyte solutions under the influence of an electric field. Analytes can be separated according to ionic mobility and / or partitioning into an alternate phase via non-covalent interactions. Additionally, analytes may be concentrated or "focused" by means of gradients in conductivity and pH.
[0028] Capillary electrophoresis separations are significant because they provide fast separations of limited sample volumes. Following reports of outstanding separation efficiencies of amines, amino acids, and peptides achieved using a 75-μm-inner-diameter glass capillary, the technology was quickly adapted to DNA.
[0029] For example, DNA samples can be incubated with a DNA-intercalating dye, followed by electrophoresis using existing polymers and buffers. The resulting electropherogram can be examined for the relative size (s) of a sample and provide a semi-quantitative measurement of the amount of DNA in the sample.
[0030] Details of capillary electrophoresis and its applications can be found, e.g. in Durney, B.C., et al. "Capillary electrophoresis applied to DNA: determining and harnessing sequence and structure to advance bioanalyses (2009–2014) . " Analytical and Bioanalytical Chemistry 407 (2015) : 6923-6938; and Kumar, R., et al. "Applications of capillary electrophoresis for biopharmaceutical product characterization. " Electrophoresis 43.1-2 (2022) : 143-166; each of which is incorporated herein by reference in its entirety.
[0031] Determination of cfDNA Concentration
[0032] In one aspect, the disclosure is related to methods for determining the concentration of cfDNA within a size range in a test sample. In some embodiments, the methods described herein include (a) isolating cfDNA from the test sample; (b) analyzing the cfDNA by capillary electropherogram to generate a distribution curve, wherein a control sample and / or a ladder are analyzed in the same batch of the cfDNA; and (c) determining the concentration of cfDNA within the size range using the ladder and / or the control sample as references. In some embodiments, the control sample comprises a pre-determined DNA concentration. In some embodiments, the ladder comprises one or more markers (e.g., DNA markers) , and at least one marker from the one or more markers has a pre-determined DNA length.
[0033] In some embodiments, the test sample described herein is blood, urine, aqueous humor, or any similar samples from a subject (e.g., a human subject) .
[0034] As generally known in the art, more than one sample can be simultaneously analyzed by capillary electropherogram, e.g., these samples can be individually loaded into the microfluidic channels on the chip for electrophoresis analysis. The microfluidic technology allows each sample to be separated and analyzed in its own channel, ensuring an accurate assessment of DNA or RNA fragment size and concentration. As described herein, these samples are regarded as being analyzed within the same batch. Results obtained from different test samples, ladders and / or control samples in the same batch (e.g., a distribution curve of the cfDNA and one or more peaks of the ladder migrated at different times, as shown in FIG. 2) can be overlaid for analysis. As discussed in Example 2, this analysis method can avoid batch-to-batch variations when results using test samples and / or ladders from different batches are compared. In addition, the sum of one or more peaks of the control sample and / or the ladder can be used as a reference to calculate the concentration of cfDNA, e.g., when the ladder markers represented by the one or more peaks are pre-determined.
[0035] In some embodiments, the ladder described herein includes a first marker and a second marker, and the size range is pre-defined by the first and second markers. In some embodiments, the lengths of the first and second markers can be within any of the size ranges described herein.
[0036] In some embodiments, the size range described herein is determined by analyzing the distribution curve of the cfDNA. In some embodiments, the size range is defined by one or more local maximum points and / or one or more local minimum points of the distribution curve. For example, the size range can be defined by one local maximum point and one local minimum point; two local maximum points; or two local minimum points. In some embodiments, the size range is defined by two neighboring local minimum points of the distribution curve. In some embodiments, the size range is defined by one local maximum point or one local minimum point, and a pre-defined ladder marker. Methods of determining local maximum points (e.g., a maxima) and local minimum points (e.g., a minima) are generally known in the art, which can be determined, e.g., by calculating the first derivative of the curve and setting it equal to zero. To differentiate the maxima versus minima, second derivatives can be calculated: a negative second derivative indicates a maxima and a positive second derivative indicates a minima. In some embodiments, local maximum points (e.g., a maxima) or local points is determined by calculating the maximum or minimum values in a pre-defined local / nearest range, which could be determined by the ladder marker.
[0037] In some embodiments, the size range described herein is about 10 bp to about 1000 bp, about 10 bp to about 900 bp, about 10 bp to about 800 bp, about 10 bp to about 700 bp, about 10 bp to about 600 bp, about 10 bp to about 500 bp, about 10 bp to about 400 bp, about 10 bp to about 300 bp, about 10 bp to about 200 bp, about 10 bp to about 100 bp, about 10 bp to about 50 bp, about 50 bp to about 1000 bp, about 50 bp to about 900 bp, about 50 bp to about 800 bp, about 50 bp to about 700 bp, about 50 bp to about 600 bp, about 50 bp to about 500 bp, about 50 bp to about 400 bp, about 50 bp to about 300 bp, about 50 bp to about 200 bp, about 50 bp to about 100 bp, about 100 bp to about 1000 bp, about 100 bp to about 900 bp, about 100 bp to about 800 bp, about 100 bp to about 700 bp, about 100 bp to about 600 bp, about 100 bp to about 500 bp, about 100 bp to about 400 bp, about 100 bp to about 300 bp, about 100 bp to about 200 bp, about 200 bp to about 1000 bp, about 200 bp to about 900 bp, about 200 bp to about 800 bp, about 200 bp to about 700 bp, about 200 bp to about 600 bp, about 200 bp to about 500 bp, about 200 bp to about 400 bp, about 200 bp to about 300 bp, about 300 bp to about 1000 bp, about 300 bp to about 900 bp, about 300 bp to about 800 bp, about 300 bp to about 700 bp, about 300 bp to about 600 bp, about 300 bp to about 500 bp, about 300 bp to about 400 bp, about 400 bp to about 1000 bp, about 400 bp to about 900 bp, about 400 bp to about 800 bp, about 400 bp to about 700 bp, about 400 bp to about 600 bp, about 400 bp to about 500 bp, about 500 bp to about 1000 bp, about 500 bp to about 900 bp, about 500 bp to about 800 bp, about 500 bp to about 700 bp, about 500 bp to about 600 bp, about 600 bp to about 1000 bp, about 600 bp to about 900 bp, about 600 bp to about 800 bp, about 600 bp to about 700 bp, about 700 bp to about 1000 bp, about 700 bp to about 900 bp, about 700 bp to about 800 bp, about 800 bp to about 1000 bp, about 800 bp to about 900 bp, or about 900 bp to about 1000 bp.
[0038] In some embodiments, the concentration of cfDNA is determined by measuring the size of an area defined by the distribution curve within the size range (e.g., any of the size ranges described herein) and a straight line connecting two local minimum points (e.g., two neighboring local minimum points) of the distribution curve. In some embodiments, the size of the area is calibrated by measuring the area size within a reference size range (e.g., the size range) from the control sample and / or the ladder across different batches. In some embodiments, the control sample and / or the ladder have the same DNA concentration among the different batches. In some embodiments, the reference size range is the same as the size range described herein. In some embodiments, the reference size range is different from the size range described herein.
[0039] In some embodiments, the size range described herein corresponds to a mono-nucleosome. In some embodiments, the size range described herein corresponds to di-nucleosome or tri-nucleosome or tetra-nucleosome.
[0040] In some embodiments, the control sample described herein includes a circulating tumor DNA reference sample prepared by inducing apoptosis in tumor cells. Details of making the reference sample can be found, e.g., in U.S. Patent Application Publication No. 20230079748A1, which is incorporated herein by reference. In some embodiments, the control sample and / or the ladder described herein include a combination of in vitro synthesized markers and those obtained from tumor cells as described above.
[0041] In some embodiments, the ladder described herein includes one or more markers (e.g., DNA markers) . In some embodiments, all or a selected group of the one or more markers have a pre-determined DNA concentration and / or length. For example, the ladder described herein may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 DNA markers. In some embodiments, the DNA markers can be in vitro synthesized. The synthesized DNA markers can have different lengths, and combined with appropriate concentrations and / or ratios to generate the ladder. For instance, the appropriate concentrations and / or ratios of the DNA markers can be determined by a skilled person in the art, such that peaks corresponding to these DNA markers do not have significant overlaps, exhibit comparable signals among themselves and relative to the target peak (e.g., the peak representing a mono-nucleosome) , and / or cover the length of the target peak in the test sample. In some embodiments, the lengths of the DNA markers are within the size range (e.g., any of the size ranges described herein) . In some embodiments, the ladder described herein includes a 150 bp marker, a200 bp marker, a250 bp marker, and / or a 300 bp marker. In some embodiments, the ladder described herein includes one or more DNA markers (represented by the one or more peaks) shown in FIG. 2 and FIG. 3.
[0042] In some embodiments, the ladder described herein includes one or more markers having a length of about 10 bp, about 20 bp, about 30 bp, about 40 bp, about 50 bp, about 60 bp, about 70 bp, about 80 bp, about 90 bp, about 100 bp, about 110 bp, about 120 bp, about 130 bp, about 140 bp, about 150 bp, about 160 bp, about 170 bp, about 180 bp, about 190 bp, about 200 bp, about 210 bp, about 220 bp, about 230 bp, about 240 bp, about 250 bp, about 260 bp, about 270 bp, about 280 bp, about 290 bp, or about 300 bp. In some embodiments, the one or more markers described herein have a length of about 10 bp to about 300 bp, about 50 bp to about 300 bp, about 100 bp to about 300 bp, about 150 bp to about 300 bp, about 200 bp to about 300 bp, about 250 bp to about 300 bp, about 10 bp to about 250 bp, about 50 bp to about 250 bp, about 100 bp to about 250 bp, about 150 bp to about 250 bp, about 200 bp to about 250 bp, about 10 bp to about 200 bp, about 50 bp to about 200 bp, about 100 bp to about 200 bp, about 150 bp to about 200 bp, about 10 bp to about 150 bp, about 50 bp to about 150 bp, about 100 bp to about 150 bp, about 10 bp to about 100 bp, about 50 bp to about 100 bp, or about 10 bp to about 50 bp. In some embodiments, the interval between any two of the markers is about 10 bp, about 20 bp, about 30 bp, about 40 bp, about 50 bp, about 60 bp, about 70 bp, about 80 bp, about 90 bp, about 100 bp, about 110 bp, about 120 bp, about 130 bp, about 140 bp, or about 150 bp.
[0043] In some embodiments, the control sample described herein includes at least one marker that has a pre-determined DNA concentration and / or length. In some embodiments, the ladder described herein includes at least one marker that has a pre-determined DNA concentration and / or length.
[0044] In some embodiments, the at least one marker has a pre-determined length of about 150 bp to about 250 bp (e.g., about 150 bp, about 160 bp, about 170 bp, about 180 bp, about 190 bp, about 200 bp, about 210 bp, about 220 bp, about 230 bp, about 240 bp, or about 250 bp) . In some embodiments, the at least one marker has a length of about 250 bp. According to NGS-based analysis results, the ending position of a mono-nucleosome is usually at about 250 bp. Thus, by overlaying results obtained from capillary electropherogram for the test sample and the control sample and / or the ladder, it is possible for a skilled person in the art to identify a peak of the distribution curve of the cfDNA that is migrated at substantially the same time or immediately before the at least one marker, and the peak represents the migrated mono-nucleosome. Thus, the methods described herein may include identifying such a peak, and the aera size under the peak can represent the concentration of cfDNA from the mono-nucleosome.
[0045] In some embodiments, the control sample and / or the ladder described herein include at least two markers (e.g., DNA markers) having pre-determined concentrations and / or lengths. In some embodiments, the lengths of the two markers are designed such that the ending position of the peak representing the mono-nucleosome is located in between. Similarly, the lengths of the two markers can be designed such that the starting position of the peak representing the mono-nucleosome is located in between. In some embodiments, the lengths of the two markers are designed such that either one or both of the starting and ending positions of the peak representing the mono-nucleosome is located in between.
[0046] In some embodiments, the ladder described herein includes a first marker having a length of about 150 bp to about 250 bp (e.g., about 150 bp, about 160 bp, about 170 bp, about 180 bp, about 190 bp, about 200 bp, about 210 bp, about 220 bp, about 230 bp, about 240 bp, or about 250 bp) , and a second marker having a length of about 250 bp to about 300 bp (e.g., about 250 bp, about 260 bp, about 270 bp, about 280 bp, about 290 bp, or about 300 bp) . In some embodiments, the first marker described herein has a length of about 150 bp to about 250 bp, about 150 bp to about 240 bp, about 150 bp to about 230 bp, about 150 bp to about 220 bp, about 150 bp to about 210 bp, about 150 bp to about 200 bp, about 150 bp to about 190 bp, about 150 bp to about 180 bp, about 150 bp to about 170 bp, about 150 bp to about 160 bp, about 160 bp to about 250 bp, about 160 bp to about 240 bp, about 160 bp to about 230 bp, about 160 bp to about 220 bp, about 160 bp to about 210 bp, about 160 bp to about 200 bp, about 160 bp to about 190 bp, about 160 bp to about 180 bp, about 160 bp to about 170 bp, about 170 bp to about 250 bp, about 170 bp to about 240 bp, about 170 bp to about 230 bp, about 170 bp to about 220 bp, about 170 bp to about 210 bp, about 170 bp to about 200 bp, about 170 bp to about 190 bp, about 170 bp to about 180 bp, about 180 bp to about 250 bp, about 180 bp to about 240 bp, about 180 bp to about 230 bp, about 180 bp to about 220 bp, about 180 bp to about 210 bp, about 180 bp to about 200 bp, about 180 bp to about 190 bp, about 190 bp to about 250 bp, about 190 bp to about 240 bp, about 190 bp to about 230 bp, about 190 bp to about 220 bp, about 190 bp to about 210 bp, about 190 bp to about 200 bp, about 200 bp to about 250 bp, about 200 bp to about 240 bp, about 200 bp to about 230 bp, about 200 bp to about 220 bp, about 200 bp to about 210 bp, about 210 bp to about 250 bp, about 210 bp to about 240 bp, about 210 bp to about 230 bp, about 210 bp to about 220 bp, about 220 bp to about 250 bp, about 220 bp to about 240 bp, about 220 bp to about 230 bp, about 230 bp to about 250 bp, about 230 bp to about 240 bp, or about 240 bp to about 250 bp. In some embodiments, the second marker described herein has a length of about 250 bp to about 300 bp, about 250 bp to about 290 bp, about 250 bp to about 280 bp, about 250 bp to about 270 bp, about 250 bp to about 260 bp, about 260 bp to about 300 bp, about 260 bp to about 290 bp, about 260 bp to about 280 bp, about 260 bp to about 270 bp, about 270 bp to about 300 bp, about 270 bp to about 290 bp, about 270 bp to about 280 bp, about 280 bp to about 300 bp, about 280 bp to about 290 bp, or about 290 bp to about 300 bp. According to the NGS-based analysis results (FIG. 1A) , the ending position of a mono-nucleosome is usually at about 250 bp. Thus, by overlaying results obtained from capillary electropherogram for the test sample and the control sample and / or the ladder, it is possible for a skilled person in the art to identify the ending position of a peak representing the migrated mono-nucleosome, by identifying the lowest point (e.g., the local minimum point) of the distribution curve of the cfDNA between the first and second markers described herein. Thus, the methods described herein may include identifying the ending position of the peak described above, and the aera size under the peak can represent the concentration of cfDNA from the mono-nucleosome.
[0047] In some embodiments, the ladder described herein includes a first marker having a length of about 10 bp to about 150 bp (e.g., about 10 bp, about 20 bp, about 30 bp, about 40 bp, about 50 bp, about 60 bp, about 70 bp, about 80 bp, about 90 bp, about 100 bp, about 110 bp, about 120 bp, about 130 bp, about 140 bp, or about 150 bp) , and a second marker having a length of about 150 bp to about 250 bp (e.g., about 150 bp, about 160 bp, about 170 bp, about 180 bp, about 190 bp, about 200 bp, about 210 bp, about 220 bp, about 230 bp, about 240 bp, or about 250 bp) . In some embodiments, the first marker described herein has a length of about 10 bp to about 150 bp, about 10 bp to about 120 bp, about 10 bp to about 100 bp, about 10 bp to about 80 bp, about 10 bp to about 60 bp, about 10 bp to about 40 bp, about 10 bp to about 20 bp, about 20 bp to about 150 bp, about 20 bp to about 120 bp, about 20 bp to about 100 bp, about 20 bp to about 80 bp, about 20 bp to about 60 bp, about 20 bp to about 40 bp, about 40 bp to about 150 bp, about 40 bp to about 120 bp, about 40 bp to about 100 bp, about 40 bp to about 80 bp, about 40 bp to about 60 bp, about 60 bp to about 150 bp, about 60 bp to about 120 bp, about 60 bp to about 100 bp, about 60 bp to about 80 bp, about 80 bp to about 150 bp, about 80 bp to about 120 bp, about 80 bp to about 100 bp, about 100 bp to about 150 bp, about 100 bp to about 120 bp, or about 120 bp to about 150 bp. In some embodiments, the second marker described herein has a length of about 150 bp to about 250 bp, about 150 bp to about 240 bp, about 150 bp to about 230 bp, about 150 bp to about 220 bp, about 150 bp to about 210 bp, about 150 bp to about 200 bp, about 150 bp to about 190 bp, about 150 bp to about 180 bp, about 150 bp to about 170 bp, about 150 bp to about 160 bp, about 160 bp to about 250 bp, about 160 bp to about 240 bp, about 160 bp to about 230 bp, about 160 bp to about 220 bp, about 160 bp to about 210 bp, about 160 bp to about 200 bp, about 160 bp to about 190 bp, about 160 bp to about 180 bp, about 160 bp to about 170 bp, about 170 bp to about 250 bp, about 170 bp to about 240 bp, about 170 bp to about 230 bp, about 170 bp to about 220 bp, about 170 bp to about 210 bp, about 170 bp to about 200 bp, about 170 bp to about 190 bp, about 170 bp to about 180 bp, about 180 bp to about 250 bp, about 180 bp to about 240 bp, about 180 bp to about 230 bp, about 180 bp to about 220 bp, about 180 bp to about 210 bp, about 180 bp to about 200 bp, about 180 bp to about 190 bp, about 190 bp to about 250 bp, about 190 bp to about 240 bp, about 190 bp to about 230 bp, about 190 bp to about 220 bp, about 190 bp to about 210 bp, about 190 bp to about 200 bp, about 200 bp to about 250 bp, about 200 bp to about 240 bp, about 200 bp to about 230 bp, about 200 bp to about 220 bp, about 200 bp to about 210 bp, about 210 bp to about 250 bp, about 210 bp to about 240 bp, about 210 bp to about 230 bp, about 210 bp to about 220 bp, about 220 bp to about 250 bp, about 220 bp to about 240 bp, about 220 bp to about 230 bp, about 230 bp to about 250 bp, about 230 bp to about 240 bp, or about 240 bp to about 250 bp. According to the NGS-based analysis results (FIG. 1A) , the starting position of a mono-nucleosome is usually at about 30-100 bp. Thus, by overlaying results obtained from capillary electropherogram for the test sample and the control sample and / or the ladder, it is possible for a skilled person in the art to identify the starting position of a peak representing the migrated mono-nucleosome, by identifying the lowest point (e.g., the local minimum point) of the distribution curve of the cfDNA between the first and second markers described herein. Thus, the methods described herein may include identifying the starting position of the peak described above, and the aera size under the peak can represent the concentration of cfDNA from the mono-nucleosome.
[0048] The methods described herein can also be optimized to determine the concentration of cfDNA within a size range corresponding to the mono-nucleosome in a more accurate fashion. As shown in FIG. 4, the distribution curve of cfDNA may exhibit a downward shift, resulting in an underestimation of the concentration of cfDNA from the mono-nucleosome. Similarly, the distribution curve of cfDNA may exhibit an upward shift, resulting in an overestimation of the concentration of cfDNA from the mono-nucleosome. To overcome such problems, the methods described herein may further include identifying the starting and / or ending positions of the peak representing the migrated mono-nucleosome described above, and the area size under the peak can be more accurately measured. For example, the concentration of cfDNA within a size range corresponding to the mono-nucleosome can be represented by the size of an area defined by the distribution curve and a straight line connecting the starting position and the ending position. In some cases, the concentration of cfDNA within a size range corresponding to the mono-nucleosome can be represented by the size of an area defined by the distribution curve within the size range (e.g., any of the size ranges described herein) and a straight line connecting two neighboring local minimum points of the distribution curve.
[0049] Because the control sample and / or the ladder described herein may include at least one marker that has a pre-determined concentration and / or length, after determining the area size under the peak representing the concentration of cfDNA from the mono-nucleosome, this area size can be calibrated by measuring the area size within a reference size range (e.g., any of the size ranges described herein) from the control sample and / or the ladder. For example, the calibration can be performed by comparing with the area size of one or more peaks of the control sample and / or the ladder (including the peak representing the at least one marker having a pre-determined DNA concentration and / or length) to determine the concentration of cfDNA from the mono-nucleosome.
[0050] In some embodiments, the method described herein further comprises determining a probability that the test sample is derived from a cancer patient based on the concentration of the cfDNA within a size range corresponding to a mono-nucleosome, determined using the methods described herein. As compared to conventional methods that correlate the total area size under the entire cfDNA distribution curve or analyzed by QubitTM fluorometer (Thermo Fisher Scientific, Q33226) , the methods described herein (e.g., CE-based methods with a control sample and / or a ladder being analyzed in the same batch) provide many benefits such as the minimal batch-to-batch variations described above. In addition, the methods described herein can lower the risk of taking into account of the contamination from longer DNA fragments, e.g., caused by hemolysis during storage or transportation, and therefore reducing the inclusion of false positive results.
[0051] As shown in FIG. 5, the cfDNA concentration within a size range corresponding to the mono-nucleosome determined using the methods described herein shows a positive correlation with the occurrence of cancer, with a p-value lower than 1×10-6. Thus, the cfDNA concentration within a size range corresponding to the mono-nucleosome can be used as a valuable indicator to determine a probability that the test sample therefrom is derived from a cancer patient, or to predict whether the subject from whom the test sample is collected from is likely to have cancer. In some embodiments, the cfDNA concentration within a size range corresponding to the mono-nucleosome of the test sample from a subject can be compared to that from a cohort of healthy individuals.
[0052] In one aspect, the disclosure is related to methods to predict cancer by determining the cfDNA concentration within a size range corresponding to the mono-nucleosome, wherein the cfDNA is isolated (e.g., extracted using any of the methods described herein) from a test sample (e.g., any of the tumor samples or healthy samples described herein) . The method can include steps of separating plasma from the sample, followed by extraction of cfDNA from the plasma, and quantifying the cfDNA concentration within a size range corresponding to the mono-nucleosome by capillary electropherogram.
[0053] In some embodiments, the cfDNA concentration within a size range corresponding to the mono-nucleosome determined from a test sample of a subject is compared with that of a reference value (e.g., the cfDNA concentration within a size range corresponding to the mono-nucleosome from a healthy subject or average cfDNA concentration within a size range corresponding to the mono-nucleosome of a group of healthy subjects) . For example, ifthe cfDNA concentration within a size range corresponding to the mono-nucleosome determined from a test sample of the subject is higher (e.g., at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 1-fold higher) than that of the reference value, the subject is likely to have cancer. In some embodiments, a ROC curve can be made according to the cfDNA concentration within a size range corresponding to the mono-nucleosome, and the AUC value can be at least or about 0.65, at least or about 0.66, at least or about 0.67, at least or about 0.68, at least or about 0.69, at least or about 0.70, at least or about 0.71, at least or about 0.72, at least or about 0.73, at least or about 0.74, at least or about 0.75, at least or about 0.76, at least or about 0.77, at least or about 0.78, at least or about 0.79, at least or about 0.80.
[0054] Determination of Short Fragment Ratio
[0055] In one aspect, the disclosure is related to methods for determining the short fragment ratio of cfDNA from a mono-nucleosome in a test sample. In some embodiments, the methods described herein include (a) isolating cfDNA from the test sample; (b) analyzing the cfDNA by capillary electropherogram to generate a distribution curve, wherein a control sample and / or a ladder are analyzed in the same batch of the cfDNA; and (c) determining the short fragment ratio of cfDNA from the mono-nucleosome using the control sample and / or the ladder as references. In some embodiments, the control sample comprises a pre-determined DNA concentration. In some embodiments, the ladder comprises one or more markers (e.g., DNA markers) , and at least one marker from the one or more markers has a pre-determined DNA length.
[0056] In some embodiments, the test sample described herein is blood, urine, aqueous humor, or any similar samples from a subject (e.g., a human subject) .
[0057] In some embodiments, results obtained from different test samples, control samples, and / or ladders in the same batch (e.g., a distribution curve of the cfDNA and one or more peaks of the ladder migrated at different times, as shown in FIG. 2) can be overlaid for analysis. Any of the control samples and / or ladders described herein can be used for determination of the short fragment ratio of cfDNA from the mono-nucleosome.
[0058] In some embodiments, the control sample described herein includes a circulating tumor DNA reference sample prepared by inducing apoptosis in tumor cells. Details of making the reference sample can be found, e.g., in U.S. Patent Application Publication No. 20230079748A1, which is incorporated herein by reference. In some embodiments, the control sample and / or the ladder described herein include a combination of in vitro synthesized markers and those obtained from tumor cells as described above.
[0059] In some embodiments, the ladder described herein includes at least one marker that has a pre-determined DNA concentration and / or length. In some embodiments, the at least one marker has a pre-determined length that is from the estimated range of the peak representing a mono-nucleosome, e.g., from about 30 bp to about 250 bp. In some embodiments, the at least one marker has a pre-determined length of about 30 bp to about 250 bp (e.g., about 30 bp, about 40 bp, about 50 bp, about 60 bp, about 70 bp, about 80 bp, about 90 bp, about 100 bp, about 110 bp, about 120 bp, about 130 bp, about 140 bp, about 150 bp, about 160 bp, about 170 bp, about 180 bp, about 190 bp, about 200 bp, about 210 bp, about 220 bp, about 230 bp, about 240 bp, or about 250 bp) . In some embodiments, the at least one marker has a length of about 30 bp to about 250 bp, about 30 bp to about 200 bp, about 30 bp to 150 bp, about 30 bp to about 100 bp, about 30 bp to about 90 bp, about 30 bp to about 80 bp, about 30 bp to about 70 bp, about 30 bp to about 60 bp, about 30 bp to about 50 bp, about 30 bp to about 40 bp, about 40 bp to about 250 bp, about 40 bp to about 200 bp, about 40 bp to 150 bp, about 40 bp to about 100 bp, about 40 bp to about 90 bp, about 40 bp to about 80 bp, about 40 bp to about 70 bp, about 40 bp to about 60 bp, about 40 bp to about 50 bp, about 50 bp to about 250 bp, about 50 bp to about 200 bp, about 50 bp to 150 bp, about 50 bp to about 100 bp, about 50 bp to about 90 bp, about 50 bp to about 80 bp, about 50 bp to about 70 bp, about 50 bp to about 60 bp, about 60 bp to about 250 bp, about 60 bp to about 200 bp, about 60 bp to 150 bp, about 60 bp to about 100 bp, about 60 bp to about 90 bp, about 60 bp to about 80 bp, about 60 bp to about 70 bp, about 70 bp to about 250 bp, about 70 bp to about 200 bp, about 70 bp to 150 bp, about 70 bp to about 100 bp, about 70 bp to about 90 bp, about 70 bp to about 80 bp, about 80 bp to about 250 bp, about 80 bp to about 200 bp, about 80 bp to 150 bp, about 80 bp to about 100 bp, about 80 bp to about 90 bp, about 90 bp to about 250 bp, about 90 bp to about 200 bp, about 90 bp to 150 bp, or about 90 bp to about 100 bp. In some embodiments, the at least one marker has a length of about 100 bp to about 250 bp, about 100 bp to about 240 bp, about 100 bp to about 230 bp, about 100 bp to about 220 bp, about 100 bp to about 210 bp, about 100 bp to about 200 bp, about 100 bp to about 190 bp, about 100 bp to about 180 bp, about 100 bp to about 170 bp, about 100 bp to about 160 bp, about 100 bp to about 150 bp, about 100 bp to about 140 bp, about 100 bp to about 130 bp, about 100 bp to about 120 bp, about 100 bp to about 110 bp, about 110 bp to about 250 bp, about 110 bp to about 240 bp, about 110 bp to about 230 bp, about 110 bp to about 220 bp, about 110 bp to about 210 bp, about 110 bp to about 200 bp, about 110 bp to about 190 bp, about 110 bp to about 180 bp, about 110 bp to about 170 bp, about 110 bp to about 160 bp, about 110 bp to about 150 bp, about 110 bp to about 140 bp, about 110 bp to about 130 bp, about 110 bp to about 120 bp, about 120 bp to about 250 bp, about 120 bp to about 240 bp, about 120 bp to about 230 bp, about 120 bp to about 220 bp, about 120 bp to about 210 bp, about 120 bp to about 200 bp, about 120 bp to about 190 bp, about 120 bp to about 180 bp, about 120 bp to about 170 bp, about 120 bp to about 160 bp, about 120 bp to about 150 bp, about 120 bp to about 140 bp, about 120 bp to about 130 bp, about 130 bp to about 250 bp, about 130 bp to about 240 bp, about 130 bp to about 230 bp, about 130 bp to about 220 bp, about 130 bp to about 210 bp, about 130 bp to about 200 bp, about 130 bp to about 190 bp, about 130 bp to about 180 bp, about 130 bp to about 170 bp, about 130 bp to about 160 bp, about 130 bp to about 150 bp, about 130 bp to about 140 bp, about 140 bp to about 250 bp, about 140 bp to about 240 bp, about 140 bp to about 230 bp, about 140 bp to about 220 bp, about 140 bp to about 210 bp, about 140 bp to about 200 bp, about 140 bp to about 190 bp, about 140 bp to about 180 bp, about 140 bp to about 170 bp, about 140 bp to about 160 bp, about 140 bp to about 150 bp, about 150 bp to about 250 bp, about 150 bp to about 240 bp, about 150 bp to about 230 bp, about 150 bp to about 220 bp, about 150 bp to about 210 bp, about 150 bp to about 200 bp, about 150 bp to about 190 bp, about 150 bp to about 180 bp, about 150 bp to about 170 bp, about 150 bp to about 160 bp, about 160 bp to about 250 bp, about 160 bp to about 240 bp, about 160 bp to about 230 bp, about 160 bp to about 220 bp, about 160 bp to about 210 bp, about 160 bp to about 200 bp, about 160 bp to about 190 bp, about 160 bp to about 180 bp, about 160 bp to about 170 bp, about 170 bp to about 250 bp, about 170 bp to about 240 bp, about 170 bp to about 230 bp, about 170 bp to about 220 bp, about 170 bp to about 210 bp, about 170 bp to about 200 bp, about 170 bp to about 190 bp, about 170 bp to about 180 bp, about 180 bp to about 250 bp, about 180 bp to about 240 bp, about 180 bp to about 230 bp, about 180 bp to about 220 bp, about 180 bp to about 210 bp, about 180 bp to about 200 bp, about 180 bp to about 190 bp, about 190 bp to about 250 bp, about 190 bp to about 240 bp, about 190 bp to about 230 bp, about 190 bp to about 220 bp, about 190 bp to about 210 bp, about 190 bp to about 200 bp, about 200 bp to about 250 bp, about 200 bp to about 240 bp, about 200 bp to about 230 bp, about 200 bp to about 220 bp, about 200 bp to about 210 bp, about 210 bp to about 250 bp, about 210 bp to about 240 bp, about 210 bp to about 230 bp, about 210 bp to about 220 bp, about 220 bp to about 250 bp, about 220 bp to about 240 bp, about 220 bp to about 230 bp, about 230 bp to about 250 bp, about 230 bp to about 240 bp, or about 240 bp to about 250 bp.
[0060] According to NGS-based analysis results (FIG. 1A) , the starting position of a mono-nucleosome is usually at about 30-100 bp, and the ending position of a mono-nucleosome is usually at about 250 bp. Thus, by overlaying results obtained from capillary electropherogram for the test sample and the control sample and / or the ladder, it is possible for a skilled person in the art to identify a peak of the distribution curve of the cfDNA that is migrated at substantially the same time with the at least one marker, and the peak represents the migrated mono-nucleosome. Thus, the methods described herein may include identifying such a peak by identifying its starting position and / or the ending position based on the distribution curve.
[0061] In some embodiments, the ladder described herein further includes at least two markers (e.g., DNA markers) having pre-determined concentrations and / or lengths. In some embodiments, the lengths of the two markers are designed such that the ending position of the peak representing the mono-nucleosome is located in between. Similarly, the lengths of the two markers can be designed such that the starting position of the peak representing the mono-nucleosome is located in between. In some embodiments, the lengths of the two markers are designed such that either one or both of the starting and ending positions of the peak representing the mono-nucleosome is located in between. By overlaying results obtained from capillary electropherogram for the test sample and the ladder, it is possible for a skilled person in the art to identify the starting and / or ending positions of a peak representing the migrated mono-nucleosome, by identifying the lowest point (e.g., the local minimum point) of the distribution curve of the cfDNA between the first and second markers (e.g., any of the first and second markers described herein) . Thus, the methods described herein may include identifying the starting and / or ending positions of the peak described above based on the distribution curve.
[0062] The methods described herein can also be optimized to determine the concentration of cfDNA from the mono-nucleosome in a more accurate fashion. As shown in FIG. 4, the distribution curve of cfDNA may exhibit a downward shift, resulting in an underestimation of the concentration of cfDNA from the mono-nucleosome. Similarly, the distribution curve of cfDNA may exhibit an upward shift, resulting in an overestimation of the concentration of cfDNA from the mono-nucleosome. To overcome such problems, the methods described herein may further include identifying the starting and / or ending positions of the peak representing the migrated mono-nucleosome described above, and the area size under the peak can be more accurately measured. For example, the concentration of cfDNA from the mono-nucleosome can be represented by the size of a first area defined by the distribution curve and a straight line connecting the starting position and the ending position. Similarly, the concentration of cfDNA shorter than the at least one marker (e.g., a marker having a pre-determined length of 150 bp) can be represented by the size of a second area defined by the distribution curve, a pre-defined ladder marker representing the at least one marker, and the straight line connecting the starting position and the ending position. The short fragment ratio of cfDNA from the mono-nucleosome described herein can be calculated by dividing the size of the second area by the size of the first area.
[0063] As another example, the concentration of cfDNA from the mono-nucleosome can be represented by the size of a first area defined by the distribution curve and a straight line connecting two neighboring local minimum points of the distribution curve of the cfDNA within the size range (e.g., any of the size ranges described herein) . In some embodiments, the concentration of short fragments from the mono-nucleosome can be represented by a second area defined by the distribution curve, a pre-defined ladder marker representing the upper size range of the short fragment (e.g., 150 bp) , and the straight line connecting the two neighboring local minimum points of the distribution curve of the cfDNA within the size range. Because the control sample and / or the ladder described herein may include at least one marker that has a pre-determined concentration and / or length, after determining the area sizes above, they can be calibrated by measuring the area size within a reference size range (e.g., any of the size ranges described herein) from the control sample and / or the ladder. For example, the calibration can be performed by comparing with the area size of one or more peaks of the control sample and / or the ladder (including the peak representing the at least one marker having a pre-determined DNA concentration and / or length) to determine the concentration of cfDNA from the mono-nucleosome. The short fragment ratio of cfDNA from the mono-nucleosome described herein can be calculated by dividing the size of the second area by the size of the first area.
[0064] In some embodiments, the method described herein further comprises determining a probability that the test sample is derived from a cancer patient based on the short fragment ratio of cfDNA from the mono-nucleosome, determined using the methods described herein. The methods described herein (e.g., CE-based methods with a control sample and / or a ladder being analyzed in the same batch) allow minimal systematic errors when the determined short fragment ratios of cfDNA from the mono-nucleosome are compared, e.g., between subjects having cancer and healthy subjects.
[0065] As shown in FIG. 7A, the short fragment ratio of cfDNA from the mono-nucleosome determined using the methods described herein shows a positive correlation with the occurrence of cancer, with a p-value lower than 1×10-5. Thus, the short fragment ratio of cfDNA from the mono-nucleosome can be used as a valuable indicator to determine a probability that the test sample therefrom is derived from a cancer patient, or to predict whether the subject from whom the test sample is collected from is likely to have cancer. In some embodiments, the short fragment ratio of cfDNA from the mono-nucleosome of the test sample from a subject can be compared to that from a cohort of heathy individuals.
[0066] In one aspect, the disclosure is related to methods to predict cancer by determining the short fragment ratio of cfDNA from the mono-nucleosome, wherein the cfDNA is isolated (e.g., extracted using any of the methods described herein) from a test sample (e.g., any of the tumor samples or healthy samples described herein) . The method can include steps of separating plasma from the sample, followed by extraction of cfDNA from the plasma, and calculating the short fragment ratio of cfDNA from the mono-nucleosome by capillary electropherogram.
[0067] In some embodiments, the short fragment ratio of cfDNA from the mono-nucleosome determined from a test sample of a subject is compared with that of a reference value (e.g., the short fragment ratio of cfDNA from the mono-nucleosome from a healthy subject or average short fragment ratio of cfDNA from the mono-nucleosome of a group of healthy subjects) . For example, if the short fragment ratio of cfDNA from the mono-nucleosome determined from a test sample of the subject is higher (e.g., at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 1-fold higher) than that of the reference value, the subject is likely to have cancer. In some embodiments, the short fragment ratio of cfDNA from the mono-nucleosome determined from a test sample of the subject (e.g., a subject having cancer) is at least 10%, at least 11%, at least 12%, at least 13%, at least 14%, at least 15%, at least 16%, at least 17%, at least 18%, at least 19%, or at least 20%. In some embodiments, the short fragment ratio of cfDNA from the mono-nucleosome determined from a test sample of the subject (e.g., a healthy subject) is less than 20%, less than 19%, less than 18%, less than 17%, less than 16%, less than 15%, less than 14%, less than 13%, less than 12%, less than 11%, or less than 10%. In some embodiments, a ROC curve can be made according to the short fragment ratio of cfDNA from the mono-nucleosome, and the AUC value can be at least or about 0.65, at least or about 0.66, at least or about 0.67, at least or about 0.68, at least or about 0.69, at least or about 0.70, at least or about 0.71, at least or about 0.72, at least or about 0.73, at least or about 0.74, at least or about 0.75, at least or about 0.76, at least or about 0.77, at least or about 0.78, at least or about 0.79, at least or about 0.80.
[0068] Measuring Markers in a Sample
[0069] Immunological measurement of blood-based PTMs has been performed over decades in clinical for cancer screening of apparently healthy individuals on large-scale; for example, alpha-fetoprotein (AFP) for liver cancer, CA125 for ovarian cancer, CA15-3 for breast cancer, CA19-9 for pancreatic cancer, CA72-4 for ovarian cancer, carcinoembryonic antigen (CEA) for cancers in digestive tract, and CYFRA 21-1 for breast carcinoma. Such methods have significant advantages including their non-invasive nature, automation, and relatively low cost compared with many other clinical detection methods (endoscopy, imaging, etc. ) . However, the low sensitivity of these methods for early cancer detection limits their widespread use for screening purposes in a general population setting.
[0070] Previous studies have shown that PTM panels are diagnostically superior to single marker for the early detection of colorectal cancer, lung cancer, breast cancer, liver cancer, gastric cancer, pancreatic cancer, ovarian cancer, and oesophagus cancer. Several reports have also demonstrated that a combined PTM panel could be used for detecting several cancer types at the same time. However, different cancer types normally show different serological characteristics. Test results can also become more complex as the sample size increases, and traditional statistical methods may not be able to handle such big data. In addition, conventional clinical methods detect multiple PTMs at the same time and use a single threshold to evaluate the results, which may cause the accumulation of false-positive rates and lead to unnecessary clinical diagnostic workups. Hence, they were not suitable for asymptomatic large-scale population screening. AI is a good analytical method for solving classification challenges by identifying implicit patterns from complex data. Over the last decade, the significant contribution of AI techniques to this advanced technology has played a critical role in medicine and healthcare research. AI is considered a valuable tool in transforming the future of healthcare and precision oncology. Several novel algorithms have shown promising results for the accurate detection and characterisation of suspected lesions.
[0071] Details of measuring markers (e.g., biomarkers) in a sample can be found, e.g., in Chinese Patent Application Publication No. 117831690A, which is incorporated herein by reference in its entirety.
[0072] Biomarkers
[0073] Before measurement can be performed a panel of markers needs to be selected for a particular cancer being screened. Many markers are known for diseases, including cancers and a known panel can be selected, or as was done as described in the examples. The panel can be selected based on measurement of individual markers in retrospective clinical samples wherein a panel is generated based on empirical data for a desired disease such as cancer, and preferably pancreatic cancer, ovarian cancer, liver cancer, lung cancer, stomach cancer, colorectal cancer, lymphoma, oesophageal cancer, or breast cancer.
[0074] Examples of biomarkers that can be employed include molecules detectable, for example, in a body fluid sample, such as, antibodies, antigens, small molecules, proteins, hormones, enzymes, genes and so on. However, the use of tumor antigens has many advantages due to their widespread use over many years and the fact that validated and standardized detection kits are available for many of them for use with the aforementioned automated immunoassay platforms.
[0075] In a particular embodiment, a panel of markers is selected based on their association with a particular cancer type. For example, AFP is a specific biomarker for liver cancer. In addition, alpha-fetoprotein (AFP) can be used as a biomarker for liver cancer (e.g., hepatocellular carcinoma) , CA125 for ovarian cancer, CA15-3 for breast cancer, CA19-9 for pancreatic cancer, CA72-4 for ovarian cancer, carcinoembryonic antigen (CEA) for cancers in digestive tract, and CYFRA 21-1 for breast carcinoma.
[0076] In certain embodiments, the panel of markers can comprise markers associated with a cancer selected from pancreatic cancer, ovarian cancer, liver cancer, lung cancer, stomach cancer, colorectal cancer, lymphoma, oesophageal cancer, prostate cancer, or breast cancer.
[0077] A panel can comprise any number of markers as a design choice, seeking, for example, to maximize specificity or sensitivity of the assay. Hence, an assay of interest may ask for presence of at least one of two or more biomarkers, three or more biomarkers, four or more biomarkers, five or more biomarkers, six or more biomarkers, seven or more biomarkers, eight or more biomarkers, nine or more biomarkers, ten or more biomarkers, or more as a design choice.
[0078] Thus, in one embodiment, the panel of biomarkers may comprise at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine or at least ten or more different markers. In one embodiment, the panel of biomarkers comprises about two to ten different markers. In another embodiment, the panel of biomarkers comprises about four to eight different markers. In yet another embodiment, the panel of markers comprises about seven different markers. In yet another embodiment, the panel of markers comprises about ten different markers.
[0079] Generally, a sample is committed to the assay and the results can be a range of numbers reflecting the presence and level (e.g., concentration, amount, activity, etc. ) of presence of each of the biomarkers of the panel in the sample.
[0080] In some embodiments, the panel of biomarkers described herein includes one or more biomarkers selected from AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1, ProGRP, SCCA, and PSA. In some embodiments, the panel of biomarkers described herein includes at least seven different biomarkers selected from AFP, CA125, CA15-3, CA19-9, CEA, CYFRA 21-1, ProGRP, SCCA, and PSA. In some embodiments, the panel of biomarkers described herein includes AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, and CYFRA 21-1. In some embodiments, the subject is a male and the panel of biomarkers described herein includes AFP, CA125, CA15-3, CA19-9, CEA, CYFRA 21-1, ProGRP, SCCA, and PSA; or in some embodiments, the subject is a female and the panel of biomarkers described herein includes AFP, CA125, CA15-3, CA19-9, CEA, CYFRA 21-1, ProGRP, and SCCA. In some embodiments, the panel of biomarkers described herein includes CEA, CYFRA21-1, SCCA, and ProGRP, and the cancer is lung cancer. In some embodiments, the panel of biomarkers described herein includes AFP, CA125, CA15-3, CA19-9, CEA, and CYFRA 21-1. In some embodiments, the subject is a male and the panel of biomarkers described herein includes AFP, CA125, CA15-3, CA19-9, CEA, CYFRA 21-1, and PSA; or in some embodiments, the subject is a female and the panel of biomarkers described herein includes AFP, CA125, CA15-3, CA19-9, CEA, and CYFRA 21-1. In some embodiments, the subject is a male and the panel of biomarkers described herein includes AFP, CA125, CA15-3, CA19-9, CEA, CYFRA 21-1, SCCA, and PSA; or in some embodiments, the subject is a female and the panel of biomarkers described herein includes AFP, CA125, CA15-3, CA19-9, CEA, CYFRA 21-1, and SCCA.
[0081] Details of biomarkers can be found, e.g., in Chinese Patent Application Publication No. 117831690A, which is incorporated herein by reference in its entirety.
[0082] Multi-Cancer Early Detection (MCED) Model
[0083] Implementations of the disclosure provide techniques for developing a machine learning system for multi-cancer early detection (MCED) tests to identify more than one type of cancer from a single test sample (e.g., a single blood sample) with high sensitivity and high accuracy. The machine learning system can be trained by one or more machine learning algorithms to distinguish cancer from non-cancer individuals.
[0084] This disclosure is based on, in part, the observational study (as shown in FIG. 8) that the two cfDNA features, including the concentration and short fragment ratio of cfDNA from a mono-nucleosome determined using the methods described herein, can complement each other in cancer detection.
[0085] In one aspect, the disclosure provides computer-implemented methods for early detection of the presence of cancer in a patient, the computer-implemented method comprising:
[0086] (a) determining the concentration of cfDNA within a size range corresponding to a mono-nucleosome in a test sample of the subject;
[0087] (b) determining the short fragment ratio of cfDNA from the mono-nucleosome in the test sample from the subject;
[0088] (c) selecting a plurality of parameters for inputs into a machine learning system, wherein the plurality of parameters comprises the concentration of cfDNA and the short fragment ratio of cfDNA from the mono-nucleosome;
[0089] (d) training the machine learning system using a machine learning algorithm selected from Random Forest (RF) , Generalized Linear Model (GLM) , Support Vector Machine (SVM) , and / or Gradient Boosting Machine (GBM) ; and
[0090] (e) determining a cancer predicting score, wherein a high cancer predicting score indicates a high probability of the subject to have cancer.
[0091] In one aspect, the disclosure provides computer-implemented methods for early detection of the presence of cancer in a patient, the computer-implemented method comprising:
[0092] (a) quantifying the level of a panel of biomarkers from a test sample (e.g., a blood sample) of the subject, wherein the panel of biomarkers comprises one or more biomarkers selected from AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA21-1, SCCA, ProGRP, and PSA;
[0093] (b) determining the concentration of cfDNA within a size range corresponding to a mono-nucleosome in the test sample of the subject;
[0094] (c) determining the short fragment ratio of cfDNA from the mono-nucleosome in the test sample of the subject;
[0095] (d) selecting a plurality of parameters for inputs into a machine learning algorithm system, wherein the plurality of parameters comprises the level of the one or more biomarkers, the concentration of cfDNA within the size range corresponding to the mono-nucleosome, and the short fragment ratio of cfDNA from the mono-nucleosome;
[0096] (e) training the machine learning system using a machine learning algorithm selected from Random Forest (RF) , Generalized Linear Model (GLM) , Support Vector Machine (SVM) , and / or Gradient Boosting Machine (GBM) ; and
[0097] (f) determining a cancer predicting score, wherein a high cancer predicting score indicates a high probability of the subject to have cancer.
[0098] In some embodiments, a generalized linear model (GLM) is used. The GLM is a flexible generalization of ordinary linear regression. The GLM generalizes linear regression by allowing the linear model to be related to the response variable via a link function and by allowing the magnitude of the variance of each measurement to be a function of its predicted value. Generalized linear models were formulated as a way of unifying various other statistical models, including linear regression, logistic regression and Poisson regression. In some embodiments, an iteratively reweighted least squares method for maximum likelihood estimation (MLE) of the model parameters is used. MLE remains popular and is the default method on many statistical computing packages. Other approaches, including Bayesian regression and least squares fitting to variance stabilized responses, have been developed.
[0099] In some embodiments, the methods described herein involve gradient boosting. Gradient boosting is a machine learning technique used in regression and classification tasks, among others. It gives a prediction model in the form of an ensemble of weak prediction models, which are typically decision trees. In some embodiments, a gradient-boosted trees model is built in a stage-wise fashion as in other boosting methods, but it generalizes the other methods by allowing optimization of an arbitrary differentiable loss function.
[0100] In some embodiments, the methods described herein involve random forest (RF) . Random forests is an ensemble learning method for classification, regression and other tasks that operates by constructing a multitude of decision trees at training time. For classification tasks, the output of the random forest is the class selected by most trees. For regression tasks, the mean or average prediction of the individual trees is returned. Random decision forests correct for decision trees'habit of overfitting to their training set.
[0101] In some embodiments, the methods described herein involve support vector machines (SVMs, also support vector networks) . SVM are supervised learning models with associated learning algorithms that analyze data for classification and regression analysis. Given a set of training examples, each marked as belonging to one of two categories, an SVM training algorithm builds a model that assigns new examples to one category or the other, making it a non-probabilistic binary linear classifier (although methods such as Platt scaling exist to use SVM in a probabilistic classification setting) . SVM maps training examples to points in space so as to maximize the width of the gap between the two categories. New examples are then mapped into that same space and predicted to belong to a category based on which side of the gap they fall. In addition to performing linear classification, SVMs can efficiently perform a non-linear classification using what is called the kernel trick, implicitly mapping their inputs into high-dimensional feature spaces.
[0102] In some embodiments, the methods described herein comprises one or more ensemble models by combining models that are trained using different algorithms, parameters, and / or training data to improve overall prediction performance. Instead of relying on a single model, the ensemble models leverage the diversity of multiple models to enhance prediction accuracy and robustness.
[0103] Different types of cancer require specific panels of PTMs, which can be used for cancer screening. However, relying solely on a single threshold for each PTM in conventional clinical methods presents challenges when combining results of mulptile PTMs, leading to an accumulation of false positives as the number of markers increases. Nevertheless, in some embodiments, the machine learning system is trained with one or more machine learning algorithms to process data for multiple cancer types together to distinguish cancer from non-cancer individuals by calculating the probability of cancer (POC) index based on the two cfDNA feature described herein (e.g., the concentration of short fragment ratio of cfDNA from a mono-nucleosome) ; the expression of multiple PTMs; and / or clinical basic information including sex and age of the individuals. By integrating multiple test results into a single one, it significantly reduces false positive rates, while maintaining the combined sensitivity of multiple tests and eliminating differences in PTMs levels among demographic groups, such as various age groups and genders. Thus, this method provides the simultaneous detection of multiple cancer types with good performance. Upon determination of whether a subject is likely to have cancer using the machine learning system, the machine learning system can be used to treat a cancer in a subject, monitor the progression of the disease, determine the effectiveness of the treatment, and adjust treatment strategy. The computer-implemented method enables early detection of cancer and more timely evaluation of treatment effectiveness, thereby reducing the likelihood of missing the optimal treatment window and lowering cancer mortality rates, ultimately improving the quality of life for cancer patients. The computer-implemented method described herein was empowered by AI technology to significantly reduce the false positive rate.
[0104] The method described herein significantly outperforms the conventional clinical method, representing a novel blood-based test for multi-cancer early detection (MCED) which is non-invasive, easy, efficient, and robust. In addition, the method described herein is affordable and accessible requiring nothing more than a blood draw at the screening sites, which makes it acceptable and sustainable in LMICs.
[0105] The present disclosure provides a protein assay that integrates the measurement of a panel (e.g., seven or ten) of selected PTMs and clinical information of the individuals, dramatically empowered by AI technology, which is more practical in LMICs.
[0106] In some embodiments, the training data set comprises at least 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10000 subjects. In some embodiments, the true positive cancer patients account for at least 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, or 80%of the sample size. In some embodiments, the training data set has no more than 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10000 subjects.
[0107] In some embodiments, the methods described herein can achieve a sensitivity of at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%. In some embodiments, the sensitivity is about 50%to about 100%, about 55%to about 100%, about 60%to about 100%, about 65%to about 100%, about 70%to about 100%, about 75%to about 100%, about 80%to about 100%, about 85%to about 100%, about 90%to about 100%, about 95%to about 100%, about 50%to about 95%, about 55%to about 95%, about 60%to about 95%, about 65%to about 95%, about 70%to about 95%, about 75%to about 95%, about 80%to about 95%, about 85%to about 95%, about 90%to about 95%, about 50%to about 90%, about 55%to about 90%, about 60%to about 90%, about 65%to about 90%, about 70%to about 90%, about 75%to about 90%, about 80%to about 90%, about 85%to about 90%, about 50%to about 85%, about 55%to about 85%, about 60%to about 85%, about 65%to about 85%, about 70%to about 85%, about 75%to about 85%, about 80%to about 85%, about 50%to about 80%, about 55%to about 80%, about 60%to about 80%, about 65%to about 80%, about 70%to about 80%, about 75%to about 80%, about 50%to about 75%, about 55%to about 75%, about 60%to about 75%, about 65%to about 75%, about 70%to about 75%, about 50%to about 70%, about 55%to about 70%, about 60%to about 70%, about 65%to about 70%, about 50%to about 65%, about 55%to about 65%, about 60%to about 65%, about 50%to about 60%, about 55%to about 60%, or about 50%to about 55%.
[0108] In some embodiments, the methods described herein can achieve a specificity of at least 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%. In some embodiments, the specificity is about 75%to about 100%, about 80%to about 100%, about 85%to about 100%, about 90%to about 100%, about 95%to about 100%, about 75%to about 95%, about 80%to about 95%, about 85%to about 95%, about 90%to about 95%, about 75%to about 90%, about 80%to about 90%, about 85%to about 90%, about 75%to about 85%, about 80%to about 85%, or about 75%to about 80%.
[0109] In some embodiments, the methods as described herein involves a cross-validation process. In some embodiments, the cross-validation process is a 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 15-fold, 20-fold, 25-fold, 30-fold, 35-fold, or 40-fold cross-validation process. In some embodiments, the cross-validation process is repeated for at least 10 times, 20 times, 30 times, 40 times, 50 times, or 100 times.
[0110] In some embodiments, the methods described herein can reduce the false-positive rate by at least 10%, 20%, 30%, 40%, or 50%as compared to a conventional method to detect cancer, e.g., a method that is based on pre-determined reference ranges for each biomarker.
[0111] In addition, in some cases, certain protein biomarkers related to multiple cancer types had a greater contribution or higher weights in the model. Conversely, certain biomarkers that were highly specific to a particular cancer type (e.g., AFP specifically for liver cancer detection) had relatively lower contributions. This led to cases where a highly specific protein biomarker for a certain cancer type exhibited abnormally high level (while others protein markers remained normal) , and the MCED model could predict a lower Probability of Cancer (POC) index.
[0112] To address this issue, an Outlier Analysis approach was developed to predict these types of cancer patients. The Outlier Analysis method focused on identifying and analyzing cases where a highly specific cancer biomarker showed exceptional expression levels compared to normal cases. By incorporating this approach into the MCED model, the detection of cancer patients who may have exhibited unique biomarker expressions was improved, providing more accurate predictions and insights for multi-cancer early cancer diagnosis. Here, we used the three methods below to determine the cutoff value for outlier analysis, based on more than 6000 non-cancer samples.
[0113] 1) Box plot method: The box plot method is used to identify outliers by plotting the protein tumor marker expression from normal control samples. A box plot displays the quartile range of the data, and observations that exceed the upper quartile plus 1.5 times the interquartile range can be considered as outliers’ cutoff value.
[0114] 2) Modified Z-Score: Since some non-cancer diseases can also result in elevated protein tumor markers expression levels in normal control cohort, the protein expression levels in the normal control cohort exhibit skewness. Therefore, the modified Z-score is considered the data skewness by calculating the difference between the observation and the median divided by the median absolute deviation (MAD) . The expression of protein tumor markers with modified Z-score>10 is defined as outliers’ cutoff value.
[0115] 3) Percentile: The percentile method compares the observation with the percentiles of the data, and observations that exceed the 99th percentile of the normal cohort can be considered as outliers’ cutoff value.
[0116] Based on the cutoff values obtained from the above three methods, the maximum value is selected as the final abnormal high outliers’ cutoff value. If the expression level of a particular biomarker in a test sample is greater than the corresponding abnormal cutoff value, the sample can be predicted as cancer patient. Through the development of this outlier analysis method, we aim to enhance the identification of cancer patients with significantly abnormal level of one cancer specific protein biomarker. By effectively predicting these exceptional cases, we can provide valuable insights into the potential presence of specific types of cancer and aid in early detection.
[0117] Thus, in some embodiments, the methods described herein further includes an Outlier Analysis. In some embodiments, the Outlier Analysis described herein involves determining the cutoff value based on at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, or 50000 non-cancer samples. In some embodiments, the cutoff value is determined by selecting the maximum value obtained from the Box plot method, Modified Z-Score, and / or Percentile described herein. In some embodiments, a cutoff value can be determined for each biomarker described herein. In some embodiments, the Outlier Analysis includes comparing the quantified level (e.g., expression level) of a selected biomarker (e.g., any of the biomarkers described herein) to its corresponding cutoff value as determined herein. For example, if the quantified level of the selected biomarker is higher than the corresponding cutoff value, there is a high probability of the patient to have cancer. In some embodiments, the Outlier Analysis is performed by (a) determining a cutoff value for each biomarker (e.g., by Box plot method, Modified Z-Score, and / or Percentile) , and (b) comparing the quantified level of each biomarker to its corresponding cutoff value. In some embodiments, a higher quantified level of a biomarker relative to its corresponding cutoff value indicates a high probability of the subject to have cancer.
[0118] Methods of Treatment
[0119] The computer-implemented method described herein can further include an additional step of treating a cancer in a subject, more timely evaluating the effectiveness of treatment, reducing the likelihood of missing the optimal treatment window, reducing the rate of the increase of volume of a tumor in a subject over time, reducing the risk of developing a metastasis, and / or reducing the risk of developing an additional metastasis in a subject, upon determination of whether the subject is likely to have cancer using the machine learning system implemented herein. In some embodiments, the treatment can halt, slow, retard, or inhibit progression of a cancer. In some embodiments, the treatment can result in the reduction of in the number, severity, and / or duration of one or more symptoms of the cancer in a subject. In some embodiments, the compositions and methods disclosed herein can be used for treatment of patients at risk for a cancer.
[0120] The treatments can generally include e.g., surgery, chemotherapy, radiation therapy, hormonal therapy, targeted therapy, and / or a combination thereof. Which treatments are used depends on the type, location and grade of the cancer as well as the patient's health and preferences. In some embodiments, the therapy is chemotherapy or chemoradiation.
[0121] In some embodiments, the disclosure is related to methods of determining whether to treat a post-surgery patient. In some embodiments, the methods comprises determining a cancer predicting score, wherein a high cancer predicting score (e.g., great than 0.8, 0.81, 0.82, 0.83, 0.84, 0.85, 0.86, 0.87, 0.88, 0.89, 0.9, 0.91, 0.92, 0.93, 0.94, or 0.95) indicates that the patient should be treated after surgery, and a low cancer predicting score (e.g., no greater than 0.8, 0.81, 0.82, 0.83, 0.84, 0.85, 0.86, 0.87, 0.88, 0.89, 0.9, 0.91, 0.92, 0.93, 0.94, or 0.95) indicates that no treatment is needed for the patient.
[0122] In one aspect, the disclosure features methods that include administering a therapeutically effective amount of a therapeutic agent to the subject in need thereof (e.g., a subject having, or identified or diagnosed as having, a cancer) . In some embodiments, the subject has e.g., breast cancer (e.g., triple-negative breast cancer) , carcinoid cancer, cervical cancer, endometrial cancer, glioma, head and neck cancer, liver cancer, lung cancer, small cell lung cancer, lymphoma, melanoma, ovarian cancer, pancreatic cancer, prostate cancer, renal cancer, colorectal cancer, gastric cancer, testicular cancer, thyroid cancer, bladder cancer, urethral cancer, or hematologic malignancy. In some embodiments, the cancer is unresectable melanoma or metastatic melanoma, non-small cell lung carcinoma (NSCLC) , small cell lung cancer (SCLC) , bladder cancer, or metastatic hormone-refractory prostate cancer. In some embodiments, the subject has a solid tumor. In some embodiments, the cancer is squamous cell carcinoma of the head and neck (SCCHN) , renal cell carcinoma (RCC) , triple-negative breast cancer (TNBC) , or colorectal carcinoma. In some embodiments, the subject has triple-negative breast cancer (TNBC) , gastric cancer, urothelial cancer, Merkel-cell carcinoma, or head and neck cancer. In some embodiments, the subject has pancreatic cancer, ovarian cancer, liver cancer, lung cancer, stomach cancer, colorectal cancer, lymphoma, oesophageal cancer, or breast cancer.
[0123] As used herein, by an “effective amount” is meant an amount or dosage sufficient to effect beneficial or desired results including halting, slowing, retarding, or inhibiting progression of a disease, e.g., a cancer. An effective amount will vary depending upon, e.g., an age and a body weight of a subject to which the therapeutic agent is to be administered, a severity of symptoms and a route of administration, and thus administration can be determined on an individual basis. An effective amount can be administered in one or more administrations. By way of example, an effective amount is an amount sufficient to ameliorate, stop, stabilize, reverse, inhibit, slow and / or delay progression of a cancer in a patient or is an amount sufficient to ameliorate, stop, stabilize, reverse, slow and / or delay proliferation of a cell (e.g., a biopsied cell, any of the cancer cells described herein, or cell line (e.g., a cancer cell line) ) in vitro.
[0124] In some embodiments, the methods described herein can be used to monitor the progression of the disease, determine the effectiveness of the treatment, and adjust treatment strategy. For example, cell free DNA (cfDNA) can be collected from the subject to detect cancer and the information can also be used to select appropriate treatment for the subject. After the subject receives a treatment, cell-free DNA can be collected from the subject. The analysis of these cfDNA can be used to monitor the progression of the disease, determine the effectiveness of the treatment, and / or adjust treatment strategy. In some embodiments, the results are then compared to the early results. In some embodiments, a dramatic increase of circulating tumor DNA indicates apoptosis at the tumor cells, which may suggest that the treatment is effective.
[0125] In some embodiments, the therapeutic agent can comprise one or more inhibitors selected from the group consisting of an inhibitor of B-Raf, an EGFR inhibitor, an inhibitor of a MEK, an inhibitor of ERK, an inhibitor of K-Ras, an inhibitor of c-Met, an inhibitor of anaplastic lymphoma kinase (ALK) , an inhibitor of a phosphatidylinositol 3-kinase (PI3K) , an inhibitor of an Akt, an inhibitor of mTOR, a dual PI3K / mTOR inhibitor, an inhibitor of Bruton's tyrosine kinase (BTK) , and an inhibitor of Isocitrate dehydrogenase 1 (IDH1) and / or Isocitrate dehydrogenase 2 (IDH2) . In some embodiments, the additional therapeutic agent is an inhibitor of indoleamine 2, 3-dioxygenase-1) (IDO1) (e.g., epacadostat) .
[0126] In some embodiments, the therapeutic agent can comprise one or more inhibitors selected from the group consisting of an inhibitor of HER3, an inhibitor of LSD1, an inhibitor of MDM2, an inhibitor of BCL2, an inhibitor of CHK1, an inhibitor of activated hedgehog signaling pathway, and an agent that selectively degrades the estrogen receptor.
[0127] In some embodiments, the therapeutic agent can comprise one or more therapeutic agents selected from the group consisting of Trabectedin, nab-paclitaxel, Trebananib, Pazopanib, Cediranib, Palbociclib, everolimus, fluoropyrimidine, IFL, regorafenib, Reolysin, Alimta, Zykadia, Sutent, temsirolimus, axitinib, everolimus, sorafenib, Votrient, Pazopanib, IMA-901, AGS-003, cabozantinib, Vinflunine, an Hsp90 inhibitor, Ad-GM-CSF, Temazolomide, IL-2, IFNa, vinblastine, Thalomid, dacarbazine, cyclophosphamide, lenalidomide, azacytidine, lenalidomide, bortezomid, amrubicine, carfilzomib, pralatrexate, and enzastaurin.
[0128] In some embodiments, the therapeutic agent can comprise one or more therapeutic agents selected from the group consisting of an adjuvant, a TLR agonist, tumor necrosis factor (TNF) alpha, IL-1, HMGB1, an IL-10 antagonist, an IL-4 antagonist, an IL-13 antagonist, an IL-17 antagonist, an HVEM antagonist, an ICOS agonist, a treatment targeting CX3CL1, a treatment targeting CXCL9, a treatment targeting CXCL10, a treatment targeting CCL5, an LFA-1 agonist, an ICAM1 agonist, and a Selectin agonist.
[0129] In some embodiments, carboplatin, nab-paclitaxel, paclitaxel, cisplatin, pemetrexed, gemcitabine, FOLFOX, or FOLFIRI are administered to the subject.
[0130] In some embodiments, the therapeutic agent is an antibody or antigen-binding fragment thereof. In some embodiments, the therapeutic agent is an antibody that specifically binds to PD-1, CTLA-4, BTLA, PD-L1, CD27, CD28, CD40, CD47, CD137, CD154, TIGIT, TIM-3, GITR, or OX40.
[0131] In some embodiments, the therapeutic agent is an anti-PD-1 antibody, an anti-OX40 antibody, an anti-PD-L1 antibody, an anti-PD-L2 antibody, an anti-LAG-3 antibody, an anti-TIGIT antibody, an anti-BTLA antibody, an anti-CTLA-4 antibody, or an anti-GITR antibody.
[0132] In some embodiments, the therapeutic agent is an anti-CTLA4 antibody (e.g., ipilimumab) , an anti-CD20 antibody (e.g., rituximab) , an anti-EGFR antibody (e.g., cetuximab) , an anti-CD319 antibody (e.g., elotuzumab) , or an anti-PD1 antibody (e.g., nivolumab) .
[0133] Systems, Software, and Interfaces
[0134] The computer-implemented methods described herein (e.g., . g., quantifying cfDNA concentration, determining cancer risk, determining the short fragment ratio, determining a cancer predicting score, etc. ) can be implemented by a computer, processor, software, module or other apparatus. Methods described herein can be computer-implemented methods, and one or more portions of a method sometimes are performed by one or more processors. Embodiments pertaining to methods described herein can be applicable to the same or related processes implemented by instructions in systems, apparatus and computer program products described herein. In some embodiments, processes and methods described herein are performed by automated methods. In some embodiments, an automated method is embodied in software, modules, processors, peripherals and / or an apparatus comprising the like, that determine sequence reads, counts, mapping, mapped sequence tags, elevations, profiles, normalizations, comparisons, range setting, categorization, adjustments, plotting, outcomes, transformations and identifications. As used herein, software refers to computer readable program instructions that, when executed by a processor, perform computer operations, as described herein.
[0135] The information for cfDNA, biomarkers, and profiles, and some other relevant information derived from a subject derived from a subject (e.g., a control subject, a patient or a subject is suspected to have tumor) can be analyzed and processed to determine the presence or absence of a genetic variation. Sequence reads and counts sometimes are referred to as “data” or “datasets” . In some embodiments, data or datasets can be characterized by one or more features or variables. In some embodiments, the sequencing apparatus is included as part of the system. In some embodiments, a system comprises a computing apparatus and a sequencing apparatus, where the sequencing apparatus is configured to receive physical nucleic acid and generate sequence reads, and the computing apparatus is configured to process the reads from the sequencing apparatus. The computing apparatus sometimes is configured to determine the presence or absence of a genetic variation (e.g., copy number variation, mutations) from the sequence reads. In some embodiments, a capillary electropherogram apparatus is included in this system.
[0136] Implementations of the subject matter and the functional operations described herein can be implemented in digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures described herein and their structural equivalents, or in combinations of one or more of the structures. Implementations of the subject matter described herein can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible program carrier for execution by, or to control the operation of, a processing device. Alternatively, or in addition, the program instructions can be encoded on a propagated signal that is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal that is generated to encode information for transmission to suitable receiver apparatus for execution by a processing device. A machine-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.
[0137] Various methods and formulae can be implemented, in the form of computer program instructions, and executed by a processing device. Suitable programming languages for expressing the program instructions include, but are not limited to, C, C++, an embodiment of FORTRAN such as FORTRAN77 or FORTRAN90, Java, Visual Basic, Perl, Tcl / Tk, JavaScript, ADA, and statistical analysis software, such as SAS, R, MATLAB, SPSS, and Stata etc. Various aspects of the computer-implemented methods may be written in different computing languages from one another, and the various aspects are caused to communicate with one another by appropriate system-level-tools available on a given system.
[0138] The processes and logic flows described in this disclosure can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input information and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit) or RISC.
[0139] Computers suitable for the execution of a computer program include, by way of example, general or special purpose microprocessors, or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and information from a read only memory or a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and information. Generally, a computer will also include, or be operatively coupled to receive information from or transfer information to, or both, one or more mass storage devices for storing information, e.g., magnetic, magneto optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a smartphone or a tablet, a touchscreen device or surface, a personal digital assistant (PDA) , a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive) , to name just a few.
[0140] Computer readable media suitable for storing computer program instructions and information include various forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto optical disks; and CD ROM and (Blue Ray) DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0141] To provide for interaction with a user, implementations of the subject matter described in this disclosure can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s client device in response to requests received from the web browser.
[0142] Implementations of the subject matter described herein can be implemented in a computing system that includes a back end component, e.g., as an information server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the subject matter, or any combination of one or more such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital information communication, e.g., a communication network. Examples of communication networks include a local area network ( “LAN” ) and a wide area network ( “WAN” ) , e.g., the Internet.
[0143] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, the server can be in the cloud via cloud computing services.
[0144] While this disclosure includes many specific implementation details, these should not be construed as limitations on the scope of any of what may be claimed, but rather as descriptions of features that may be specific to particular implementations. Certain features that are described in this disclosure in the context of separate implementations can also be implemented in combination in a single implementation. Conversely, various features that are described in the context of a single implementation can also be implemented in multiple implementations separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
[0145] Similarly, while operations are described in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system components in the implementations described above should not be understood as requiring such separation in all implementations, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0146] Particular implementations of the subject matter have been described. Other implementations are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. In one embodiment, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some implementations, multitasking and parallel processing may be advantageous.
[0147] In one aspect, the disclosure provides various computer-implemented methods as described herein.
[0148] In one aspect, the disclosure provides one or more machine-readable hardware storage devices for the computer-implemented methods as described herein by storing instructions that are executable by one or more data processing devices to perform operations as described herein.
[0149] In one aspect, the disclosure provides a system comprising: one or more data processing devices; and one or more machine-readable hardware storage devices for the computer-implemented methods for a test subject by storing instructions that are executable by the one or more data processing devices to perform operations as described herein.
[0150] Kits
[0151] The present disclosure also provides kits for collecting, transporting, and / or analyzing samples. Such a kit can include materials and reagents required for obtaining an appropriate sample (e.g., cfDNA) from a subject, or for measuring the levels of particular biomarkers. In some embodiments, the kits include those materials and reagents that would be required for obtaining and storing a sample from a subject. The sample is then shipped to a service center for further processing (e.g., sequencing and / or data analysis) .
[0152] The kits may further include instructions for collect the samples, performing the assay and methods for interpreting and analyzing the data resulting from the performance of the assay.
[0153] EXAMPLES
[0154] The invention is further described in the following examples, which do not limit the scope of the invention described in the claims.
[0155] EXAMPLE 1: Calculation of Short Fragment Ratio Based on Capillary Electropherogram
[0156] In recent years, many studies have reported that the ratio of short fragments of cell-free DNA (cfDNA) in blood / urine samples from cancer patients is significantly higher than that in blood / urine samples from healthy individuals. Almost all current methods used to detect the ratio of short fragments rely on the Next-Generation Sequencing (NGS) approach, where fragment size is determined by aligning paired-end (PE) sequencing reads to the human genome and calculating their positional differences. This wet process is complex and costly. Here, a new method is developed by calculating the short fragment ratio using capillary electropherogram.
[0157] According to Meng, Z., et al. "Noninvasive detection of hepatocellular carcinoma with circulating tumor DNA features and α-fetoprotein. " The Journal of Molecular Diagnostics 23.9 (2021) : 1174-1184) , the short fragment ratio of a sample, e.g., the proportion of cfDNA fragments with fragment size (FS) less than 150 bp (P150) , is defined as the number of fragments shorter than 150 bp divided by the total cfDNA fragment number derived from a mono-nucleosome. As shown in FIG. 1A, the short fragment ratio, e.g., P150, can be determined by dividing the number of reads with the fragment size less than 150 bp by the total reads with fragment size within the size of a mono-nucleosome, according to the NGS sequencing results.
[0158] As described herein, the short fragment ratio can also be calculated by determining the starting and ending positions of the mono-nucleosome, as well as the position corresponding to 150 bp using capillary electropherogram (FIG. 1B) . Based on these three positions, the concentration of fragments shorter than 150 bp (the area under the curve from the starting position to the 150 bp mark) and the concentration of cfDNA fragments from the mono-nucleosome (the area under the curve from the starting position to the ending position of the mono-nucleosome) were calculated. The ratio between these two concentrations was equal to the short fragment ratio (e.g., P150) .
[0159] Methods
[0160] Sample collection and analysis
[0161] Blood collection, plasma separation, and cfDNA extraction from the subjects can be performed using standard cfDNA extraction procedures. Detailed processes can be found in Example 1 of Chinese Patent No. 112397143B, which is incorporated herein by reference in its entirety.
[0162] Capillary electrophoresis analysis of the extracted cfDNA was performed using an Agilent 2100 Bioanalyzer. Specifically, 1μL of cfDNA was used with the Agilent 2100 Bioanalyzer (Agilent, Model: G29939BA) in combination with the Agilent High Sensitivity DNA Kit (Agilent, Cat#: 5067-4626) to detect the cfDNA peaks and determine the concentrations (as reflected by absorbance) at different fragment lengths (which were migrated at different times) .
[0163] Determination of starting position, ending position, and 150 bp position
[0164] We overlaid the ladder results and the sample results in a single graph. As shown in FIG. 2, the black line represents the result of one sample's electropherogram, while the grey line represents the ladder result. The ladder was designed with a 150 bp marker (peak) , and the corresponding migration time of this peak was used as the P150 cutoff value for all samples processed in the same batch. Thus, the area under the curve before this position represents the concentration of short fragments (e.g., those less than 150 bp) . As another example, ifthere is a need to calculate the concentration of fragments between 90 bp and 150 bp, a corresponding ladder can be prepared based on the methods described above by determining the migration times corresponding to the 90 bp peak and 150 bp peak.
[0165] To determine the starting and ending positions of the mono-nucleosome, the ending position was defined based on the mono-nucleosome range specified in the NGS results from Meng, Z., et al. as described above (e.g., 250 bp) . A250 bp ladder marker can be designed and prepared using the methods described above. This marker was used to directly determine the migration time of the 250 bp peak in the same batch of samples, such that the concentration of the mono-nucleosome can be calculated.
[0166] In addition, a new method was developed to determine the starting and ending positions of the mono-nucleosome, which can be directly predicted based on the curve of cfDNA fragment size distribution. According to Zhu, D., et al. "Circulating cell-free DNA fragmentation is a stepwise and conserved process linked to apoptosis. " BMC Biology 21.1 (2023) : 253, when the migrated number of cfDNA fragments of a mono-nucleosome approaches zero in the fragment size (FS) distribution of the entire cfDNA sequencing results, the starting and ending positions of the mono-nucleosome can be predicted by determining the lowest points of the distribution curve. As shown in FIG. 1A, these lowest points are shown as the troughs on both sides of the mono-nucleosome peak at around 167 bp. For example, because the lowest point between the main peak of the first mono-nucleosome (167 bp) and the main peak of the second mono-nucleosome (333 bp) is located at around 250 bp, the ending position of the first mono-nucleosome is expected to be located within this range. Therefore, the ending position was determined by identifying the lowest point between the 150 bp and 300 bp ladder markers in the sample's electropherogram (FIG. 2) . Similarly, the lowest point before 150 bp was defined as the starting position of the mono-nucleosome.
[0167] EXAMPLE 2: Calculation of cfDNA Concentration Within the Mono-Nucleosome Range Based on Capillary Electropherogram
[0168] By comparing experimental results using different batches, significant batch-to-batch variations were observed. As shown in FIG. 3, although the same ladder was used in different batches, the migration times (X-axis) of peaks having the same cfDNA fragment sizes differed between batches. Because of this, it was contemplated that using fixed migration times to determine the starting position, ending position, and 150 bp position may be problematic. Rather, using ladders from the same batch to determine the 150 bp position and using the lowest points within a certain range to define the starting and ending positions, may be necessary.
[0169] It was also observed that the absorbance fluorescence intensity (Y-axis) for a certain peak in the same sample differed in different batches. To address this problem, concentration calibration was performed. Specifically, the sum of the peak values of ladder peaks (e.g., 50 bp, 100 bp, 150 bp, 200 bp, and 300 bp peaks) was used as a reference for batch-to-batch calibration. Alternatively, a standard reference that mimics cfDNA distribution can be prepared, as detailed in the U.S. Patent Application Publication No. 20230079748A1, which is incorporated herein by reference in its entirety. The mono-nucleosome concentration of this standard reference can be used for batch-to-batch calibration.
[0170] Calculation of the corrected cfDNA concentration from a mono-nucleosome
[0171] Due to various factors, the fragment size (FS) distribution curve of samples often shows upward or downward shifts. In a normal sample (FIG. 2) , the starting and ending positions of the mono-nucleosome nearly coincide with the baseline (y=0) . The NGS sequencing graph as shown in FIG. 1A also showed this. However, if the curve shifts downwards, most of the area of the mono-nucleosome is below the baseline (y=0) , resulting in a significantly underestimated mono-nucleosome concentration when calculated directly (FIG. 4) .
[0172] To address this problem, a fitted line based on the starting and ending positions of the mono-nucleosome can be drawn (see the dashed line in FIG. 4) . Next, the final cfDNA concentration from the mono-nucleosome can be calculated as the area between the mono-nucleosome curve and this fitted straight line. Similarly, an upward shift of the curve may also be observed. In both cases, one can first determine the starting and ending positions of the mono-nucleosome according to the methods described above, and then draw a fitted straight line based on the determined starting and ending positions to calculate the corrected concentration.
[0173] Optimized cfDNA concentration from a mono-nucleosome for multi-cancer early detection (MCED)
[0174] According to Chinese Patent No. 112397143B, there is a significant difference of the cfDNA concentration in the plasma between non-cancer individuals and cancer patients. Here, we compared the cfDNA concentration from a mono-nucleosome determined by the conventional methods (v0) and the corrected cfDNA concentration determined by the optimized methods discussed herein (v2) in a cohort (including 71 cancer samples and 120 non-cancer samples) .
[0175] The conventional cfDNA concentration may be calculated by the following methods. For example, when the Agilent 2100 Bioanalyzer is used, the cfDNA concentration is determined by the total area under the entire FS distribution curve. Alternatively, the cfDNA concentration can be quantified using a QubitTM fluorometer (Thermo Fisher Scientific, Q33226) , as detailed in Chinese Patent No. 112397143B. Here, the second method above (using the QubitTM fluorometer) was used as a reference (cfDNA v0) for comparison with the optimized methods discussed herein (cfDNA v2) .
[0176] In conventional methods, some samples may experience hemolysis due to delayed plasma separation after blood collection or improper transportation temperatures. For example, when blood cells break down, their intracellular DNA may enter the plasma, causing genomic DNA (gDNA) contamination. This may also lead to an overall increase of cfDNA concentration. In Chinese Patent No. 112397143B, it was found that cfDNA concentrations can be significantly higher in tumor samples as compared to that in normal samples, which makes the detection results prone to include false positives. For the fragment size of gDNA is always more than 1,000bp. Therefore, the cfDNA concentration from the mono-nucleosome was calculated based on the methods described above, instead of using the total cfDNA concentration, to lower the risk of taking account of the contamination from longer DNA fragments. This advantage may help eliminating false positives caused by hemolyzed samples.
[0177] As shown in FIG. 5, in a retrospective cohort, the optimized cfDNA concentration from the mono-nucleosome (cfDNA v2) showed a significantly higher value in cancer patients, compared to the healthy controls (p-value=3.4×10-7) . As shown in FIG. 6, the area under curve (AUC) value of cfDNA v2 was 0.721. Compared to the conventional method (using the total area under the entire FS distribution curve, cfDNA v0) , the AUC increased by 0.063 when the optimized cfDNA concentration was used. Notably, at high specificity levels such as 90%, the sensitivity improved by about 9.6%.
[0178] EXAMPLE 3: Short Fragment Ratio Based on Capillary Electropherogram For Cancer Early Detection
[0179] Methods of calculating the cfDNA concentration of fragment size smaller than 150 bp is discussed in Example 2, e.g., by drawing a fitted straight line between the starting and ending positions of the mono-nucleosome. Specifically, after determining the intersection point of the straight line with the 150 bp mark, one can draw a line from this intersection point to the starting point. The correction methods described above can be used to calculate the final area to obtain the concentration of short fragments smaller than 150 bp. Finally, one can divide the cfDNA concentration of fragments smaller than 150 bp by the cfDNA concentration from the mono-nucleosome to obtain the short fragment ratio (e.g., P150) .
[0180] The same cohort as described in Example 2 was used to calculate the short fragment ratio, which showed a significant difference between cancer and non-cancer (healthy) samples, with a p-value=6×10-4 by t-test (FIG. 7A) . The AUC value for P150 v2, or the short fragment ratio obtained using the optimized methods discussed herein (cfDNA v2) , was 0.745 (FIG. 7B) .
[0181] EXAMPLE 4: Combining Capillary Electropherogram Features Using Machine Learning for MCED
[0182] Using a 90%specificity threshold to detect cancer patients, we compared the true positive cases identified by cfDNA concentration from a mono-nucleosome (Example 2, FIG. 6) with those identified by the short fragment ratio (Example 3, FIG. 7B) . As shown in FIG. 8, these two methods were found to complement each other in cancer detection. Therefore, a machine learning model, such as Generalized Linear Model (GLM) , Random Forest (RF) , or Gradient Boosting Machine (GBM) , was employed to combine the short fragment ratio with cfDNA concentration from a mono-nucleosome to predict cancer probability. Specifically, a Random Forest (RF) model was used for this purpose.
[0183] The model can be built as follows. First, using the methods described in Examples 2-3, the short fragment ratio and the cfDNA concentration from a mono-nucleosome for each sample were obtained. The short fragment ratio and cfDNA concentration from the mono-nucleosome for all samples were normalized using the median and MAD (median absolute deviation) values from healthy samples. The Modified Z-Score was calculated by subtracting the healthy median from the sample value, then dividing by the MAD to account for variability. This approach provides a robust measure of how far each sample deviates from the healthy baseline. Second, the standardized short fragment ratio and cfDNA concentration from the mono-nucleosome were used as features to train a Random Forest (RF) model by ten-fold cross-validation. The average value predicted from these models is defined as the final cancer probability of the test sample.
[0184] Using the same cohort as described in Example 2 the probability of predicting cancer showed a highly significant difference between cancer and non-cancer samples, with p-value=3.5×10-12 by t-test (FIG. 9A) . The AUC value for the probability of predicting cancer was 0.782 (FIG. 9B) .
[0185] EXAMPLE 5: Combining Multidimensional Cancer Features Using Machine Learning for MCED
[0186] In past years, clinical practice has utilized methods that detect one or more protein biomarkers in blood for cancer screening, such as PSA for prostate cancer screening, AFP for hepatocellular carcinoma screening, CA125 for ovarian cancer screening, CA153 for breast cancer screening, CA199 for pancreatic cancer screening, CA724 for ovarian cancer screening, CEA for gastrointestinal cancer screening, and CYFRA21-1 for breast cancer screening. Compared to other clinical tests, these methods offer advantages such as being non-invasive, automated, and cost-effective. These methods can also be used for early screening of multiple cancers and cancer tracing. More details can be found, e.g., in Chinese Patent Application Publication No. 117831690A, which is incorporated herein by reference in its entirety.
[0187] By incorporating protein biomarkers in blood as features into the model obtained from Example 4, improved cancer screening performance can be achieved. Additionally, both methods (with or without protein biomarkers in blood incorporated into the model) do not involve NGS, which can significantly reduce the testing cost.
[0188] To build the model based on that described in Example 4, seven additional protein biomarkers (AFP, CA125, CA153, CA199, CA724, CEA, and CYFRA21-1) , along with the two features from the capillary electropherogram (cfDNA concentration from the mono-nucleosome and short fragment ratio) were used as inputs to build the cancer detection model. According to the steps described in Example 4, the RF method was used to combine these standardized features to build the model and to calculate the probability of predicting cancer for the test samples. As shown in FIG. 10A, the probability of predicting cancer showed a highly significant difference between cancer and non-cancer samples, with a p-value=4.67×10-22 by t-test. The AUC value for the probability of predicting cancer was 0.890 (FIG. 10B) .
[0189] OTHER EMBODIMENTS
[0190] It is to be understood that while the invention has been described in conjunction with the detailed description thereof, the foregoing description is intended to illustrate and not limit the scope of the invention, which is defined by the scope of the appended claims. Other aspects, advantages, and modifications are within the scope of the following claims.
Claims
1.A method for determining the concentration of cell-free DNA (cfDNA) within a size range in a test sample, the method comprising:(a) isolating cfDNA from the test sample;(b) analyzing the cfDNA by capillary electropherogram to generate a distribution curve, wherein a ladder and / or a control sample are analyzed in the same batch of the cfDNA, wherein the control sample comprises a pre-determined DNA concentration, wherein the ladder comprises one or more markers, and at least one marker from the one or more markers has a pre-determined DNA length; and(c) determining the concentration of the cfDNA within the size range using the control sample and / or the ladder as references.2.The method of claim 1, wherein the control sample is a circulating tumor DNA reference sample prepared by inducing apoptosis in tumor cells.3.The method of claim 1 or 2, wherein the distribution curve exhibits an upward or downward shift.4.The method of any one of claims 1-3, wherein the ladder comprises a first marker and a second marker, and the size range is pre-defined by the first marker and the second marker.5.The method of any one of claims 1-3, wherein the size range is determined by analyzing the distribution curve of the cfDNA.6.The method of claim 5, wherein the size range is defined by one or more local maximum points and / or one or more local minimum points of the distribution curve (e.g., one local maximum point and a local minimum point of the distribution curve) .7.The method of claim 5 or 6, wherein the size range is defined by two neighboring local minimum points of the distribution curve.8.The method of claim 5 or 6, wherein the size range is defined by one local maximum point or one local minimum point, and a pre-defined ladder marker.9.The method of claim 7, wherein the concentration of cfDNA is determined by measuring the size of an area defined by the distribution curve within the size range and a straight line connecting the two neighboring local minimum points of the distribution curve, and optionally the size of the area is calibrated by measuring the area size within a reference size range (e.g., the size range) from the control sample and / or the ladder across different batches, wherein the control sample and / or the ladder have the same DNA concentration among the different batches.10.The method of any one of claims 1-9, wherein the size range corresponds to a mono-nucleosome.11.A method for determining the short fragment ratio of cfDNA from a mono-nucleosome in a test sample, the method comprising:(a) isolating cfDNA from the test sample;(b) analyzing the cfDNA by capillary electropherogram to generate a distribution curve, wherein a control sample and / or a ladder are analyzed in the same batch of the cfDNA, wherein the control sample comprises a pre-determined DNA concentration, wherein the ladder comprises one or more markers, and at least one marker from the one or more markers has a pre-determined DNA length; and(c) determining the short fragment ratio of cfDNA from the mono-nucleosome using the control sample and / or the ladder as references.12.The method of claim 11, wherein the control sample is a circulating tumor DNA reference sample prepared by inducing apoptosis in tumor cells.13.The method of claim 11 or 12, wherein the distribution curve exhibits an upward or downward shift.14.The method of any one of claims 11-13, wherein step (c) comprisesidentifying a size range of the cfDNA that corresponds to the mono-nucleosome;determining a first area size defined by the distribution curve and a straight line connecting two neighboring local minimum points of the distribution curve of the cfDNA within the size range;determining a second area size defined by the distribution curve, a pre-defined ladder marker representing the upper size range of the short fragment (e.g., 150 bp) , and the straight line connecting the two neighboring local minimum points of the distribution curve of the cfDNA within the size range, and optionally the first and second area sizes are calibrated by measuring the area size within a reference size range (e.g., the size range of the cfDNA that corresponds to the mono-nucleosome) from the control sample and / or the ladder across different batches, wherein the control sample and / or the ladder have the same DNA concentration among the different batches; andcalculating the short fragment ratio of cfDNA from the mono-nucleosome by dividing the second area size by the first area size.15.The method of any one of claims 1-14, wherein the method further comprises determining a probability that the test sample is derived from a cancer patient based on the concentration of the cfDNA within a size range that corresponds to a mono-nucleosome and / or the short fragment ratio of cfDNA from the mono-nucleosome.16.A computer-implemented method for early detection of the presence of cancer in a subject, the method comprising:(a) determining the concentration of cfDNA within a size range corresponding to a mono-nucleosome in a test sample of the subject;(b) determining the short fragment ratio of cfDNA from the mono-nucleosome in the test sample from the subject;(c) selecting a plurality of parameters for inputs into a machine learning system, wherein the plurality of parameters comprises the concentration of cfDNA and the short fragment ratio of cfDNA from the mono-nucleosome;(d) training the machine learning system using a machine learning algorithm selected from Random Forest (RF) , Generalized Linear Model (GLM) , Support Vector Machine (SVM) , and / or Gradient Boosting Machine (GBM) ; and(e) determining a cancer predicting score, wherein a high cancer predicting score indicates a high probability of the subject to have cancer.17.A computer-implemented method for early detection of the presence of cancer in a subject, the method comprising:(a) quantifying the level of a panel of biomarkers from a test sample (e.g., a blood sample) of the subject, wherein the panel of biomarkers comprises one or more protein biomarkers selected from AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA21-1, SCCA, ProGRP, and PSA;(b) determining the concentration of cfDNA within a size range corresponding to a mono-nucleosome in the test sample of the subject;(c) determining the short fragment ratio of cfDNA from the mono-nucleosome in the test sample of the subject;(d) selecting a plurality of parameters for inputs into a machine learning algorithm system, wherein the plurality of parameters comprises the level of the one or more protein biomarkers, the concentration of cfDNA within the size range corresponding to the mono-nucleosome, and the short fragment ratio of cfDNA from the mono-nucleosome;(e) training the machine learning system using a machine learning algorithm selected from Random Forest (RF) , Generalized Linear Model (GLM) , Support Vector Machine (SVM) , and / or Gradient Boosting Machine (GBM) ; and(f) determining a cancer predicting score, wherein a high cancer predicting score indicates a high probability of the subject to have cancer.18.The method of claim 17, wherein the plurality of parameters further comprises at least one clinical parameter (e.g., age, gender, and / or smoking status) .19.The method of any one of claims 16-18, wherein the concentration of cfDNA within the size range corresponding to the mono-nucleosome and the short fragment ratio of cfDNA from the mono-nucleosome are determined by capillary electropherogram.20.The method of any one of claims 16-19, wherein the concentration of cfDNA within the size range corresponding to the mono-nucleosome and the short fragment ratio of cfDNA from the mono-nucleosome are determined byisolating cfDNA from the test sample;analyzing the cfDNA by capillary electropherogram to generate a distribution curve, wherein a control sample and / or a ladder are analyzed in the same batch of the cfDNA; anddetermining the concentration of cfDNA within the size range corresponding to the mono-nucleosome and the short fragment ratio of cfDNA from the mono-nucleosome using the control sample and / or the ladder as references.21.The method of any one of claims 16-20, wherein the method can aid early detection of the presence of at least two cancer types simultaneously.
Citation Information
Patent Citations
USing cell-free DNA fragment size to determine copy number variations
CN107750277A
Using cell-free DNA fragment size to detect tumor-associated variant
CN110800063A
Cell-free DNA end characteristics
CN113366122A
Detection of lung cancer using cell free DNA fragmentation
CN116940994A
Computer implementation method for detecting abnormal signal quantification of blood sample to be detected
CN117831690A