Cancer risk prediction method based on capillary electrophoresis fragment characteristics of cell-free DNA

By analyzing the concentration of cfDNA and the proportion of short mononuclear body fragments in plasma using capillary electrophoresis, combined with protein biomarkers and artificial intelligence, the accuracy and cost issues of existing cancer screening methods have been resolved, enabling more efficient cancer detection and early warning.

CN122459673APending Publication Date: 2026-07-24H & H WORKS LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
H & H WORKS LTD
Filing Date
2025-04-27
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing cancer screening methods are inadequate in terms of accuracy and cost, especially in low- or middle-income countries, and struggle to achieve early detection and efficient screening of multiple cancer types.

Method used

By analyzing the concentration of cfDNA in plasma and the proportion of short fragments from mononuclear bodies using capillary electrophoresis, combined with protein biomarkers and artificial intelligence, a more sensitive, accurate, and cost-effective cancer detection method can be developed.

Benefits of technology

This technology enables more accurate prediction of whether a sample comes from a cancer patient at a low cost, making it suitable for early detection and recurrence monitoring. It reduces detection costs and improves the sensitivity and specificity of the test.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122459673A_ABST
    Figure CN122459673A_ABST
Patent Text Reader

Abstract

The present invention relates to methods for determining the concentration of cell-free DNA (cfDNA) from capillary electropherogram and calculating the proportion of short fragments from mono-nucleosomes. In some embodiments, the determination of cfDNA concentration and mono-nucleosome short fragment proportion, optionally in combination with a panel of screened protein biomarkers, can be used to achieve efficient and cost-effective early detection of multiple cancer types.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for determining the concentration of cell-free DNA (cfDNA) and the proportion of short fragments derived from mononuclear bodies using capillary electrophoresis. Specifically, this application relates to a method, system, electronic device, and computer-readable medium for predicting the probability that a test sample originates from a cancer patient based on plasma characteristics and artificial intelligence. Background Technology

[0002] Numerous studies have shown that circulating tumor DNA (ctDNA) fragments from tumor cells in the blood are shorter than normal cell-free DNA (cfDNA), and the size of cfDNA fragments can be assessed using next-generation sequencing (NGS). Furthermore, significant differences exist in cfDNA fragmentation patterns in the genome between healthy subjects and cancer patients, as well as between different cancer types. Current standard clinical practice (SOC) cancer screening methods, including imaging, plasma tumor markers, and cytology, are limited to specific cancer types and suffer from unsatisfactory accuracy and participant compliance. Moreover, current methods heavily rely on NGS to detect cfDNA fragmentation patterns, which is both complex and expensive. Therefore, there is a need to develop more efficient and cost-effective methods for early cancer detection, recurrence monitoring, treatment response assessment, and mechanistic studies of individual cancer etiologies. Summary of the Invention

[0003] This invention addresses one of the technical problems in related fields. In this regard, it provides a non-invasive method for cancer detection, recurrence monitoring, and treatment response assessment based on multidimensional features of cell-free DNA (cfDNA) and / or protein biomarkers in plasma, combined with artificial intelligence. This method is based on a cancer genome panorama combined with protein biomarkers. The technology interprets plasma characteristics using capillary electrophoresis. Simultaneously, by combining specific protein biomarkers with big data and artificial intelligence, it can predict the probability that a test sample originates from a cancer patient. Based on multiple features of the test sample (including cfDNA concentration and the proportion of short fragments from mononuclear bodies), this invention employs a multidimensional, multivariate weighted algorithm, combining cfDNA features (e.g., based on cfDNA concentration and the proportion of short fragments from mononuclear bodies) and / or protein biomarkers (e.g., based on the level of a set of protein biomarkers in the blood), thereby predicting the probability that a test sample originates from a cancer patient in a more sensitive, accurate, and specific manner while keeping detection costs more manageable. Compared to NGS-based techniques, the detection method described herein is less complex and more cost-effective. In one aspect, the present invention relates to a method for determining the concentration of cell-free DNA (cfDNA) within a specific size range in a test sample, the method comprising: (a) isolating cfDNA from the test sample; (b) analyzing the cfDNA by capillary electrophoresis to generate a distribution curve, wherein in some embodiments, a molecular weight standard and / or a control sample are analyzed in the same batch as the cfDNA, wherein in some embodiments, the control sample contains a predetermined DNA concentration, wherein in some embodiments, the molecular weight standard comprises one or more molecular weight markers, and at least one of the molecular weight markers has a predetermined DNA length; and (c) using the control sample and / or the molecular weight standard as a reference, determining the concentration of cfDNA within the size range. In some embodiments, the control sample is a circulating tumor DNA reference sample prepared by inducing tumor cell apoptosis. In some embodiments, the distribution curve exhibits an upward or downward shift. In some embodiments, the molecular weight standard includes a first molecular weight marker and a second molecular weight marker, and the size range is predefined by the first molecular weight marker and the second molecular weight marker. In some embodiments, the size range is determined by analyzing the distribution curve of the cfDNA. In some embodiments, the size range is defined by one or more local maxima and / or one or more local minima of the distribution curve (e.g., a local maximum and a local minimum of the distribution curve). In some embodiments, the size range is defined by two adjacent local minima of the distribution curve.In some embodiments, the size range is defined by a local maximum or a local minimum and a predefined molecular weight standard. In some embodiments, the concentration of cfDNA is determined by measuring the area defined by the distribution curve within the size range and the straight line connecting two adjacent local minimums of the distribution curve, and optionally, the area is calibrated by measuring the area of ​​control samples and / or molecular weight standards across different batches within a reference size range (e.g., the size range), in some embodiments where the control samples and / or molecular weight standards have the same DNA concentration across different batches. In some embodiments, the size range corresponds to a mononucleosome.

[0004] It should be noted that the analytical molecular weight standard (ladder) mentioned above, along with the reference molecular weight, molecular weight standard, or length standard, all refer to a reference standard composed of nucleic acid fragments of known length. It is used to calibrate the size of nucleic acid fragments during capillary electrophoresis to establish the correspondence between migration time and fragment length, thereby achieving accurate determination and quality control of sample fragment size distribution.

[0005] In one aspect, the present invention relates to a method for determining the proportion of short cfDNA fragments from mononuclear bodies in a test sample, the method comprising: (a) isolating cfDNA from the test sample; (b) analyzing the cfDNA by capillary electrophoresis to generate a distribution curve, wherein in some embodiments, a control sample and / or a molecular weight standard are analyzed in the same batch as the cfDNA, wherein in some embodiments, the control sample contains a predetermined DNA concentration, and in some embodiments, the molecular weight standard comprises one or more molecular weight standard points, and at least one of the one or more molecular weight standard points has a predetermined DNA length; and (c) using the control sample and / or the molecular weight standard as a reference, determining the proportion of short cfDNA fragments from mononuclear bodies. In some embodiments, the control sample is a circulating tumor DNA reference sample prepared by inducing tumor cell apoptosis. In some embodiments, the distribution curve exhibits an upward or downward shift. In some embodiments, step (c) includes: identifying a size range of cfDNA corresponding to the mononuclear body; determining a first area size defined by a distribution curve of cfDNA within the size range and a straight line connecting two adjacent local minima of the distribution curve; determining a second area size defined by the distribution curve, a predetermined molecular weight standard point representing the upper limit size range of short fragments (e.g., 150 bp), and a straight line connecting two adjacent local minima of the cfDNA distribution curve within the size range, and optionally, calibrating the first and second area sizes by measuring the area sizes of control samples and / or molecular weight standards across different batches within a reference size range (e.g., the cfDNA size range corresponding to the mononuclear body), wherein in some embodiments, the control samples and / or molecular weight standards have the same DNA concentration across different batches; and calculating the proportion of cfDNA short fragments from the mononuclear body by dividing the second area size by the first area size. In some embodiments, the method further includes: determining the probability that the test sample originated from a cancer patient based on the cfDNA concentration corresponding to the size range of the mononuclear body and / or the proportion of cfDNA short fragments from the mononuclear body.In one aspect, the present invention relates to a computer-implemented method for early detection of the presence of cancer in a subject, the method comprising: (a) determining the concentration of cfDNA in a test sample of the subject within a size range corresponding to a single nucleosome; (b) determining the proportion of short cfDNA fragments from single nucleosomes in the test sample of the subject; (c) selecting multiple parameters to input into a machine learning system, wherein in some embodiments, the multiple parameters include cfDNA concentration and the proportion of short cfDNA fragments from single nucleosomes; (d) training the machine learning system using a machine learning algorithm selected from random forest (RF), generalized linear model (GLM), support vector machine (SVM), and / or gradient boosting machine (GBM); and (e) determining a cancer prediction score, wherein in some embodiments, a higher cancer prediction score indicates a higher probability that the subject has cancer. In one aspect, the present invention relates to a computer-implemented method for early detection of the presence of cancer in a subject, the method comprising: (a) quantifying the levels of a set of biomarkers in a subject test sample (e.g., a blood sample), wherein in some embodiments, the set of biomarkers comprises one or more protein biomarkers selected from AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA21-1, SCCA, ProGRP, and PSA; (b) determining the concentration of cfDNA in the subject test sample corresponding to a single nucleosome size range; and (c) determining the concentration of cfDNA in the subject test sample. (d) Selecting multiple parameters to input into a machine learning system, in some embodiments, the multiple parameters including the level of one or more protein biomarkers, the concentration of cfDNA corresponding to the size range of the mononucleus, and the proportion of cfDNA short fragments from the mononucleus; (e) training the machine learning system using a machine learning algorithm selected from random forest (RF), generalized linear model (GLM), support vector machine (SVM), and / or gradient boosting machine (GBM); and (f) determining a cancer prediction score, in some embodiments, a higher cancer prediction score indicating a higher probability that the subject has cancer. In some embodiments, the multiple parameters also include at least one clinical parameter (e.g., age, sex, and / or smoking status). In some embodiments, the concentration of cfDNA corresponding to the size range of the mononucleus and the proportion of cfDNA short fragments from the mononucleus are determined by capillary electrophoresis.In some embodiments, the concentration of cfDNA corresponding to the size range of mononuclear bodies and the proportion of short cfDNA fragments derived from mononuclear bodies are determined by the following steps: isolating cfDNA from a test sample; analyzing the cfDNA by capillary electrophoresis to generate a distribution curve; in some embodiments, analyzing a control sample and / or molecular weight standard in the same batch as the cfDNA; and using the control sample and / or the molecular weight standard as a reference, determining the concentration of cfDNA corresponding to the size range of mononuclear bodies and the proportion of short cfDNA fragments derived from mononuclear bodies. In some embodiments, the method can assist in the simultaneous early detection of the presence of at least two cancer types.

[0006] It should be noted that the above-mentioned SCCA, SCC, or SCC antigen all refer to squamous cell carcinoma antigen.

[0007] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. This document describes the methods and materials used in this invention; other suitable methods and materials known in the art may also be used. These materials, methods, and examples are illustrative only and are not intended to be limiting. All publications, patent applications, patents, sequences, database entries, and other references mentioned herein are incorporated herein by reference in their entirety. In the event of any conflict, this specification (including definitions) shall prevail.

[0008] Other features and advantages of the invention will be apparent from the following detailed description, drawings, and claims. Attached Figure Description

[0009] Figure 1 A shows the distribution of cfDNA fragment sizes based on next-generation sequencing (NGS). Figure 1 B shows the cfDNA fragment size distribution based on capillary electrophoresis. This result was obtained using an Agilent 2100 Bioanalyzer. The start, 150 bp, and end positions of the mononuclear bodies are marked. The Y-axis represents the fluorescence intensity of the labeled nucleic acids, which is correlated with their concentration in the sample. “FU” represents fluorescence units. The X-axis represents migration time.

[0010] Figure 2This image shows an overlaid profile of the electrophoresis results (black line) and the molecular weight standard (ladder) results (gray line) of a sample. The results were obtained using Agilent 2100 bioanalyzer software (Agilent Technologies Inc.). The 150 bp peak (“150 bp”), 200 bp peak (“200 bp”), and 300 bp peak (“300 bp”) of the molecular weight standard are marked. The lowest point between the 150 bp and 300 bp peaks represents the termination position of the mononuclear body. The Y-axis represents peak fluorescence intensity, and the X-axis represents migration time.

[0011] Figure 3 The superimposed spectra of molecular weight standards from different batches are displayed. The Y-axis represents peak fluorescence intensity, and the X-axis represents migration time. The right side shows magnified results of peaks that migrated before 90 seconds. The horizontal and vertical offsets demonstrate the necessity of inter-batch calibration.

[0012] Figure 4 This displays an electrophoretic plot of a sample with a downward offset. The fitted straight line (shown as a dashed line) can be used as a baseline for calculating the area under the curve.

[0013] Figure 5 Box plot results show the optimized cfDNA concentrations from mononuclei in cancer patients (“cancer patients”) and healthy individuals (“healthy individuals”).

[0014] Figure 6 ROC (Receiver Operating Characteristic) curves indicating the performance of cancer detection using different cfDNA concentration assay methods are shown. The optimized cfDNA concentration from mononuclear bodies (cfDNA v2) shows superior performance compared to the conventional method (cfDNA v0).

[0015] Figure 7 A shows the box plot results of the proportion of short cfDNA fragments (i.e. fragments less than 150 bp) in cancer patients (“cancer patients”) and healthy individuals (“healthy people”). Figure 7 B shows the ROC curve indicating cancer detection performance based on the proportion of short cfDNA fragments (i.e., fragments less than 150 bp), which was determined using optimized cfDNA concentrations from mononuclei (P150 v2).

[0016] Figure 8 This diagram shows the complementary relationship between the concentration of cfDNA from mononuclear bodies (cfDNA v2) and the proportion of short fragments (P150 v2) when predicting true positive samples. The numbers in the Venn diagram represent the number of true positive cases detected by cfDNA v2, P150 v2, or both.

[0017] Figure 9 A shows box plot results predicting the probability of cancer in cancer patients (“cancer patients”) and healthy individuals (“healthy individuals”), obtained using a machine learning model that combines cfDNA concentration (cfDNA v2) and short fragment ratio (P150 v2) from mononuclear bodies. Figure 9 B shows the use of Figure 9 The ROC curve of the machine learning model in A, which predicts the probability of cancer to distinguish between cancer patients and non-cancer individuals.

[0018] Figure 10 A shows the box plot results predicting the probability of cancer in cancer patients (“cancer patients”) and healthy individuals (“healthy individuals”), obtained using a machine learning model that incorporates cfDNA concentration from mononuclear bodies (cfDNA v2), the proportion of short fragments (P150 v2), and seven protein biomarkers (AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA21-1). Figure 10 B shows the use of Figure 10 The ROC curve of the machine learning model in A, which predicts the probability of cancer to distinguish between cancer patients and non-cancer individuals. Detailed Implementation

[0019] Cancer is a significant public health problem worldwide. The global cancer burden is rapidly increasing, with an estimated 19.3 million new cases and 10 million cancer deaths in 2020. It is estimated that more than two-thirds of the world's annual cancer deaths occur in low- or middle-income countries (LMICs). The global cancer burden is projected to reach 28.4 million cases by 2040, a 47% increase from 2020. Due to demographic shifts, the increase will be greater in transition countries (64% to 95%) than in established transition countries (32% to 56%), although this may be further exacerbated by increasing risk factors associated with limited healthcare infrastructure in LMICs. Early cancer detection is well-known for providing higher cure rates and 5-year survival rates, and for reducing treatment costs and economic productivity losses. Various cancer screening technologies are currently available in clinical practice. Examples include low-dose computed tomography (LDCT) for lung cancer screening, mammography for breast cancer detection, HPV testing or cytology combined with colposcopy for early cervical cancer detection, fecal occult blood test (FOBT) combined with colonoscopy for colorectal cancer screening, and prostate-specific antigen (PSA) for prostate cancer. However, the high cost of these screening methods and their reliance on specialized infrastructure and skilled personnel limit their application, which is why alternative screening tests exist in LMICs: for example, visual acetic acid staining (VIA) for cervical cancer and clinical breast examination (CBE) for breast cancer screening. Furthermore, these methods are each designed to screen for specific cancer types, hindering their widespread use as universal screening tools. Furthermore, liquid biopsy methods for detecting analytes in blood, such as cancer-derived DNA, are currently being used in human medicine to screen for multiple types of cancer simultaneously; these MCED (Multi-Cancer Early Detection) tests represent a paradigm shift in cancer screening and promise to significantly increase the number of cancer patients detected at earlier stages. However, due to cost, complexity, and reliance on high-end infrastructure and rigorous laboratories, these tests are not suitable for use in LMICs (Limited Microorganisms). Taken together, these factors contribute to cancer often being diagnosed at a late stage and exacerbate healthcare inequalities. Developing and validating a more universal, robust, and cost-effective MCED test is crucial for enabling large-scale cancer screening in seemingly healthy individuals, particularly in LMIC populations, in the future. This invention discloses a method for determining the concentration of cell-free DNA (cfDNA) via capillary electrophoresis and a method for determining the proportion of short fragments from mononuclear bodies.These two cfDNA features, optionally combined with the levels of a selected set of protein biomarkers in the blood, can be used as input to train machine learning systems to predict cancer in a more sensitive, accurate, and cost-effective manner.

[0020] Capillary electrophoresis and its application in DNA fragment analysis Capillary electrophoresis (CE) is a series of electrokinetic separation methods performed in sub-millimeter diameter capillaries, as well as microfluidic and nanofluidic channels. Typically, CE refers to capillary zone electrophoresis (CZE), but other electrophoretic techniques, including capillary gel electrophoresis (CGE), capillary isoelectric focusing (CIEF), capillary isovelocity electrophoresis, and micellar electrokinetic chromatography (MEKC), also fall into this category. In CE methods, analytes migrate through an electrolyte solution under the influence of an electric field. Analytes can be separated based on ion mobility and / or partitioning into an alternative phase via non-covalent interactions. Furthermore, analytes can be concentrated or “focused” using gradients in conductivity and pH.

[0021] Capillary electrophoresis is significant because it provides rapid separation for limited sample volumes. Following reports of excellent separation efficiency for amines, amino acids, and peptides using glass capillaries with an inner diameter of 75 micrometers, this technique has been rapidly applied to DNA analysis.

[0022] For example, Durney, BC et al., "Capillary electrophoresis applied to DNA: determining and harnessing sequence and structure to advance bioanalyses (2009–2014)" Analytical and Bioanalytical Chemistry 407 (2015): 6923-6938; and Kumar, R. et al., "Applications of capillary electrophoresis for biopharmaceutical product characterization." Electrophoresis 43.1-2 (2022): 143-166; each of the above references is incorporated herein by reference in its entirety.

[0023] Determination of cfDNA concentration In one aspect, the present invention relates to a method for determining the concentration of cfDNA within a specific size range in a test sample. In some embodiments, the method described herein includes: (a) isolating cfDNA from the test sample; (b) analyzing the cfDNA by capillary electrophoresis to generate a distribution curve, wherein a control sample and / or a molecular weight standard (ladder) are analyzed in the same batch as the cfDNA; and (c) determining the concentration of cfDNA within the size range using the molecular weight standard and / or the control sample as a reference. In some embodiments, the control sample contains a predetermined DNA concentration. In some embodiments, the molecular weight standard comprises one or more molecular weight standard spots (e.g., DNA molecular weight standard spots), and at least one of the one or more molecular weight standard spots has a predetermined DNA length. In some embodiments, the test sample described herein is blood, urine, aqueous humor, or any similar sample from a subject (e.g., a human subject). As is generally known in the art, multiple samples can be analyzed simultaneously by capillary electrophoresis, for example, these samples can be individually loaded into microfluidic channels on a chip for electrophoretic analysis. Microfluidic technology allows each sample to be separated and analyzed in its own channel, thereby ensuring accurate assessment of the size and concentration of DNA or RNA fragments. As described herein, these samples are considered to be analyzed within the same batch. Results obtained from different test samples, molecular weight standards, and / or control samples within the same batch (e.g., cfDNA distribution curves and so on) can be compared. Figure 2 The peaks of one or more molecular weight standards that migrate at different times are superimposed for analysis. As discussed in Example 2, this analytical method avoids batch-to-batch differences when comparing results using test samples and / or molecular weight standards from different batches. Furthermore, the sum of one or more peaks of the control sample and / or molecular weight standards can be used as a reference for calculating cfDNA concentration, for example, when the molecular weight standard points represented by the one or more peaks are predetermined.

[0024] In some embodiments, the molecular weight standards described herein include a first molecular weight standard point and a second molecular weight standard point, and the size range is predefined by the first and second molecular weight standard points. In some embodiments, the lengths of the first and second molecular weight standard points may be within any size range described herein.

[0025] In some embodiments, the size range described herein is determined by analyzing the distribution curve of cfDNA. In some embodiments, the size range is defined by one or more local maxima and / or one or more local minima of the distribution curve. For example, the size range may be defined by one local maximum and one local minimum; two local maxima; or two local minima. In some embodiments, the size range is defined by two adjacent local minima of the distribution curve. In some embodiments, the size range is defined by one local maximum or one local minimum and a predefined molecular weight standard point. Methods for determining local maxima (e.g., maxima) and local minima (e.g., minima) are commonly known in the art; for example, they can be determined by calculating the first derivative of the curve and setting it equal to zero. To distinguish between maxima and minima, the second derivative can be calculated: a negative second derivative indicates a maxima, and a positive second derivative indicates a minima. In some embodiments, a local maximum (e.g., a maximum) or a local minimum (e.g., a minimum) is determined by calculating the maximum or minimum value within a predetermined local / nearest range, which may be determined by a molecular weight standard point.

[0026] In some embodiments, the size range described herein is approximately 10bp to approximately 1000bp, approximately 10bp to approximately 900bp, approximately 10bp to approximately 800bp, approximately 10bp to approximately 700bp, approximately 10bp to approximately 600bp, approximately 10bp to approximately 500bp, approximately 10bp to approximately 400bp, approximately 10bp to approximately 300bp, approximately 10bp to approximately 200bp, approximately 10bp to approximately 100bp, approximately 10bp to approximately 50bp, approximately 50bp to approximately 1000bp, approximately 50bp to approximately 900bp, approximately 50bp to approximately 800bp, approximately 50bp to approximately 700bp, approximately 50bp to approximately 600bp, approximately 50bp to... Approximately 500bp, approximately 50bp to approximately 400bp, approximately 50bp to approximately 300bp, approximately 50bp to approximately 200bp, approximately 50bp to approximately 100bp, approximately 100bp to approximately 1000bp, approximately 100bp to approximately 900bp, approximately 100bp to approximately 800bp, approximately 100bp to approximately 700bp, approximately 100bp to approximately 600bp, approximately 100bp to approximately 500bp, approximately 100bp to approximately 400bp, approximately 100bp to approximately 300bp, approximately 100bp to approximately 200bp, approximately 200bp to approximately 1000bp, approximately 200bp to approximately 900bp, approximately 200bp to approximately 800bp, approximately 200 bp to 700bp, approximately 200bp to 600bp, approximately 200bp to 500bp, approximately 200bp to 400bp, approximately 200bp to 300bp, approximately 300bp to 1000bp, approximately 300bp to 900bp, approximately 300bp to 800bp, approximately 300bp to 700bp, approximately 300bp to 600bp, approximately 300bp to 500bp, approximately 300bp to 400bp, approximately 400bp to 1000bp, approximately 400bp to 900bp, approximately 400bp to 800bp, approximately 400bp to 700bp, approximately 400bp to 600bp bp, approximately 400bp to approximately 500bp, approximately 500bp to approximately 1000bp, approximately 500bp to approximately 900bp, approximately 500bp to approximately 800bp, approximately 500bp to approximately 700bp, approximately 500bp to approximately 600bp, approximately 600bp to approximately 1000bp, approximately 600bp to approximately 900bp, approximately 600bp to approximately 800bp, approximately 600bp to approximately 700bp, approximately 700bp to approximately 1000bp, approximately 700bp to approximately 900bp, approximately 700bp to approximately 800bp, approximately 800bp to approximately 1000bp, approximately 800bp to approximately 900bp, or approximately 900bp to approximately 1000bp.

[0027] In some embodiments, the concentration of cfDNA is determined by measuring the area defined by a distribution curve within the size range (e.g., any size range described herein) and a straight line connecting two local minima (e.g., two adjacent local minima) of the distribution curve. In some embodiments, the area is calibrated by measuring the area of ​​control samples and / or molecular weight standards across different batches within a reference size range (e.g., the size range). In some embodiments, the control samples and / or molecular weight standards have the same DNA concentration across different batches. In some embodiments, the reference size range is the same as the size range described herein. In some embodiments, the reference size range differs from the size range described herein.

[0028] In some embodiments, the size range described herein corresponds to a mono-nucleosome. In some embodiments, the size range described herein corresponds to a di-nucleosome, a tri-nucleosome, or a tetra-nucleosome.

[0029] In some embodiments, the control sample described herein includes a circulating tumor DNA reference sample prepared by inducing tumor cell apoptosis. Details regarding the preparation of this reference sample can be found, for example, in U.S. Patent Application Publication No. 20230079748A1, which is incorporated herein by reference. In some embodiments, the control sample and / or molecular weight standard (ladder) described herein includes a combination of in vitro synthesized molecular weight standard points and molecular weight standard points obtained from tumor cells as described above.

[0030] In some embodiments, the molecular weight standards described herein include one or more molecular weight standard points (e.g., DNA molecular weight standard points). In some embodiments, all or a selected group of the one or more molecular weight standard points have a predetermined DNA concentration and / or length. For example, the molecular weight standards described herein may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 DNA molecular weight standard points. In some embodiments, the DNA molecular weight standard points may be synthesized in vitro. Synthesized DNA molecular weight standard points may have different lengths and be combined in appropriate concentrations and / or proportions to generate molecular weight standards. For example, those skilled in the art can determine appropriate concentrations and / or proportions of DNA molecular weight standard points such that the peaks corresponding to these DNA molecular weight standard points do not exhibit significant overlap, exhibit comparable signals among themselves and relative to the target peak (e.g., a peak representing a mononuclear body), and / or cover the length of the target peak in the test sample. In some embodiments, the length of the DNA molecular weight standard spot is within the size range (e.g., any size range described herein). In some embodiments, the molecular weight standards described herein include 150 bp, 200 bp, 250 bp, and / or 300 bp molecular weight standard spots. In some embodiments, the molecular weight standards described herein include, for example... Figure 2 and Figure 3 One or more DNA molecular weight standard points are shown (represented by one or more peaks).

[0031] In some embodiments, the molecular weight standards described herein include one or more molecular weight standard points with lengths of about 10 bp, about 20 bp, about 30 bp, about 40 bp, about 50 bp, about 60 bp, about 70 bp, about 80 bp, about 90 bp, about 100 bp, about 110 bp, about 120 bp, about 130 bp, about 140 bp, about 150 bp, about 160 bp, about 170 bp, about 180 bp, about 190 bp, about 200 bp, about 210 bp, about 220 bp, about 230 bp, about 240 bp, about 250 bp, about 260 bp, about 270 bp, about 280 bp, about 290 bp, or about 300 bp. In some embodiments, the lengths of one or more molecular weight standard points described herein are approximately 10 bp to approximately 300 bp, approximately 50 bp to approximately 300 bp, approximately 100 bp to approximately 300 bp, approximately 150 bp to approximately 300 bp, approximately 200 bp to approximately 300 bp, approximately 250 bp to approximately 300 bp, approximately 10 bp to approximately 250 bp, approximately 50 bp to approximately 250 bp, approximately 100 bp to approximately 250 bp, approximately 150 bp to approximately 250 bp, approximately 200 bp to approximately 250 bp, approximately 10 bp to approximately 200 bp, approximately 50 bp to approximately 200 bp, approximately 100 bp to approximately 200 bp, approximately 150 bp to approximately 200 bp, approximately 10 bp to approximately 150 bp, approximately 50 bp to approximately 150 bp, approximately 10 bp to approximately 150 bp, approximately 10 bp to approximately 100 bp, approximately 50 bp to approximately 100 bp, or approximately 10 bp to approximately 50 bp. In some embodiments, the interval between any two molecular weight standard points is about 10 bp, about 20 bp, about 30 bp, about 40 bp, about 50 bp, about 60 bp, about 70 bp, about 80 bp, about 90 bp, about 100 bp, about 110 bp, about 120 bp, about 130 bp, about 140 bp, or about 150 bp.

[0032] In some embodiments, the control samples described herein include at least one molecular weight standard point having a predetermined DNA concentration and / or length. In some embodiments, the molecular weight standard described herein includes at least one molecular weight standard point having a predetermined DNA concentration and / or length.

[0033] In some embodiments, the at least one molecular weight standard spot has a predetermined length of about 150 bp to about 250 bp (e.g., about 150 bp, about 160 bp, about 170 bp, about 180 bp, about 190 bp, about 200 bp, about 210 bp, about 220 bp, about 230 bp, about 240 bp, or about 250 bp). In some embodiments, the length of the at least one molecular weight standard spot is about 250 bp. According to NGS-based analysis results, the termination position of a mononuclear body is typically at about 250 bp. Therefore, by superimposing the results obtained from capillary electrophoresis patterns of test samples and control samples and / or molecular weight standards, those skilled in the art can identify a peak in the cfDNA distribution curve that migrates at or immediately before the at least one molecular weight standard spot, and that the peak represents a migrating mononuclear body. Therefore, the methods described herein may include identifying such a peak, and the area under the peak may represent the concentration of cfDNA from mononuclear bodies.

[0034] In some embodiments, the control samples and / or molecular weight standards described herein include at least two molecular weight standard spots (e.g., DNA molecular weight standard spots) having predetermined concentrations and / or lengths. In some embodiments, the lengths of the two molecular weight standard spots are designed such that the termination position of the peak representing a mononuclear body is located between them. Similarly, the lengths of the two molecular weight standard spots may be designed such that the start position of the peak representing a mononuclear body is located between them. In some embodiments, the lengths of the two molecular weight standard spots are designed such that either or both of the start and termination positions of the peak representing a mononuclear body are located between them.

[0035] In some embodiments, the molecular weight standards described herein include a first molecular weight standard point of about 150 bp to about 250 bp (e.g., about 150 bp, about 160 bp, about 170 bp, about 180 bp, about 190 bp, about 200 bp, about 210 bp, about 220 bp, about 230 bp, about 240 bp, or about 250 bp) and a second molecular weight standard point of about 250 bp to about 300 bp (e.g., about 250 bp, about 260 bp, about 270 bp, about 280 bp, about 290 bp, or about 300 bp). In some embodiments, the length of the first molecular weight standard point described herein is approximately 150 bp to approximately 250 bp, approximately 150 bp to approximately 240 bp, approximately 150 bp to approximately 230 bp, approximately 150 bp to approximately 220 bp, approximately 150 bp to approximately 210 bp, approximately 150 bp to approximately 200 bp, approximately 150 bp to approximately 190 bp, approximately 150 bp to approximately 180 bp, approximately 150 bp to approximately 170 bp, approximately 150 bp to approximately 160 bp, approximately 160 bp to approximately 250 bp, approximately 160 bp to approximately 240 bp, approximately 160 bp Approximately 160bp to approximately 220bp, approximately 160bp to approximately 210bp, approximately 160bp to approximately 200bp, approximately 160bp to approximately 190bp, approximately 160bp to approximately 180bp, approximately 160bp to approximately 170bp, approximately 170bp to approximately 250bp, approximately 170bp to approximately 240bp, approximately 170bp to approximately 230bp, approximately 170bp to approximately 220bp, approximately 170bp to approximately 210bp, approximately 170bp to approximately 200bp, approximately 170bp to approximately 190bp, approximately 170bp to approximately 1 80bp, approximately 180bp to approximately 250bp, approximately 180bp to approximately 240bp, approximately 180bp to approximately 230bp, approximately 180bp to approximately 220bp, approximately 180bp to approximately 210bp, approximately 180bp to approximately 200bp, approximately 180bp to approximately 190bp, approximately 190bp to approximately 250bp, approximately 190bp to approximately 240bp, approximately 190bp to approximately 230bp, approximately 190bp to approximately 220bp, approximately 190bp to approximately 210bp, approximately 190bp to approximately 200bp, approximately 200bp to approximately 250bp p, approximately 200bp to approximately 240bp, approximately 200bp to approximately 230bp, approximately 200bp to approximately 220bp, approximately 200bp to approximately 210bp, approximately 210bp to approximately 250bp, approximately 210bp to approximately 240bp, approximately 210bp to approximately 230bp, approximately 210bp to approximately 220bp, approximately 220bp to approximately 250bp, approximately 220bp to approximately 240bp, approximately 220bp to approximately 230bp, approximately 230bp to approximately 250bp, approximately 230bp to approximately 240bp, or approximately 240bp to approximately 250bp.In some embodiments, the length of the second molecular weight standard point described herein is about 250 bp to about 300 bp, about 250 bp to about 290 bp, about 250 bp to about 280 bp, about 250 bp to about 270 bp, about 250 bp to about 260 bp, about 260 bp to about 300 bp, about 260 bp to about 290 bp, about 260 bp to about 280 bp, about 260 bp to about 270 bp, about 270 bp to about 300 bp, about 270 bp to about 290 bp, about 270 bp to about 280 bp, about 280 bp to about 300 bp, about 280 bp to about 290 bp, or about 290 bp to about 300 bp.

[0036] Based on NGS-based analysis results ( Figure 1 A) The termination position of a mononuclear body is typically at approximately 250 bp. Therefore, by overlaying the results obtained from capillary electrophoresis plots of the test sample and control samples and / or molecular weight standards, those skilled in the art can identify the termination position of a peak representing a migrating mononuclear body by recognizing the lowest point (e.g., a local minimum) of the cfDNA distribution curve between the first and second molecular weight standard points described herein. Thus, the methods described herein may include identifying the termination position of the aforementioned peak, and the area under that peak may represent the concentration of cfDNA from the mononuclear body.

[0037] In some embodiments, the molecular weight standards described herein include a first molecular weight standard point of about 10 bp to about 150 bp (e.g., about 10 bp, about 20 bp, about 30 bp, about 40 bp, about 50 bp, about 60 bp, about 70 bp, about 80 bp, about 90 bp, about 100 bp, about 110 bp, about 120 bp, about 130 bp, about 140 bp, or about 150 bp) and a second molecular weight standard point of about 150 bp to about 250 bp (e.g., about 150 bp, about 160 bp, about 170 bp, about 180 bp, about 190 bp, about 200 bp, about 210 bp, about 220 bp, about 230 bp, about 240 bp, or about 250 bp). In some embodiments, the length of the first molecular weight standard point described herein is approximately 10 bp to approximately 150 bp, approximately 10 bp to approximately 120 bp, approximately 10 bp to approximately 100 bp, approximately 10 bp to approximately 80 bp, approximately 10 bp to approximately 60 bp, approximately 10 bp to approximately 40 bp, approximately 10 bp to approximately 20 bp, approximately 20 bp to approximately 150 bp, approximately 20 bp to approximately 120 bp, approximately 20 bp to approximately 100 bp, approximately 20 bp to approximately 80 bp, approximately 20 bp to approximately 60 bp, approximately 20 bp to approximately 40 bp, approximately 40 bp p to about 150bp, about 40bp to about 120bp, about 40bp to about 100bp, about 40bp to about 80bp, about 40bp to about 60bp, about 60bp to about 150bp, about 60bp to about 120bp, about 60bp to about 100bp, about 60bp to about 80bp, about 80bp to about 150bp, about 80bp to about 120bp, about 80bp to about 100bp, about 100bp to about 150bp, about 100bp to about 120bp, or about 120bp to about 150bp.In some embodiments, the length of the second molecular weight standard point described herein is approximately 150 bp to approximately 250 bp, approximately 150 bp to approximately 240 bp, approximately 150 bp to approximately 230 bp, approximately 150 bp to approximately 220 bp, approximately 150 bp to approximately 210 bp, approximately 150 bp to approximately 200 bp, approximately 150 bp to approximately 190 bp, approximately 150 bp to approximately 180 bp, approximately 150 bp to approximately 170 bp, approximately 150 bp to approximately 160 bp, approximately 160 bp to approximately 250 bp, approximately 160 bp to approximately 240 bp, approximately 160 bp Approximately 160bp to approximately 220bp, approximately 160bp to approximately 210bp, approximately 160bp to approximately 200bp, approximately 160bp to approximately 190bp, approximately 160bp to approximately 180bp, approximately 160bp to approximately 170bp, approximately 170bp to approximately 250bp, approximately 170bp to approximately 240bp, approximately 170bp to approximately 230bp, approximately 170bp to approximately 220bp, approximately 170bp to approximately 210bp, approximately 170bp to approximately 200bp, approximately 170bp to approximately 190bp, approximately 170bp to approximately 1 80bp, approximately 180bp to approximately 250bp, approximately 180bp to approximately 240bp, approximately 180bp to approximately 230bp, approximately 180bp to approximately 220bp, approximately 180bp to approximately 210bp, approximately 180bp to approximately 200bp, approximately 180bp to approximately 190bp, approximately 190bp to approximately 250bp, approximately 190bp to approximately 240bp, approximately 190bp to approximately 230bp, approximately 190bp to approximately 220bp, approximately 190bp to approximately 210bp, approximately 190bp to approximately 200bp, approximately 200bp to approximately 250bp p, approximately 200bp to approximately 240bp, approximately 200bp to approximately 230bp, approximately 200bp to approximately 220bp, approximately 200bp to approximately 210bp, approximately 210bp to approximately 250bp, approximately 210bp to approximately 240bp, approximately 210bp to approximately 230bp, approximately 210bp to approximately 220bp, approximately 220bp to approximately 250bp, approximately 220bp to approximately 240bp, approximately 220bp to approximately 230bp, approximately 230bp to approximately 250bp, approximately 230bp to approximately 240bp, or approximately 240bp to approximately 250bp. Based on NGS-based analysis results (…). Figure 1A) The initiation position of mononuclear bodies is typically around 30-100 bp. Therefore, by overlaying the results obtained from capillary electrophoresis plots of test samples and control samples and / or molecular weight standards, those skilled in the art can identify the initiation position of a peak representing a migrating mononuclear body by recognizing the lowest point (e.g., a local minimum) of the cfDNA distribution curve between the first and second molecular weight standard points described herein. Therefore, the methods described herein may include identifying the initiation position of the aforementioned peak, and the area under that peak may represent the concentration of cfDNA from the mononuclear body.

[0038] The method described in this article can be further optimized to more accurately determine cfDNA concentrations within a size range corresponding to single nucleosomes. For example... Figure 4 As shown, the cfDNA distribution curve may be downward shifted, leading to an underestimation of the cfDNA concentration from mononuclear bodies. Similarly, the cfDNA distribution curve may be upward shifted, leading to an overestimation of the cfDNA concentration from mononuclear bodies. To overcome these problems, the method described herein may further include identifying the start and / or end positions of the peak representing the aforementioned migration of mononuclear bodies, and can more accurately measure the area under that peak. For example, the cfDNA concentration corresponding to a mononuclear body size range may be represented by the area defined by the distribution curve and the straight line connecting the start and end positions. In some cases, the cfDNA concentration corresponding to a mononuclear body size range may be represented by the area defined by the distribution curve within said size range (e.g., any size range described herein) and the straight line connecting two adjacent local minima of said distribution curve.

[0039] Because the control samples and / or molecular weight standards described herein may include at least one molecular weight standard spot having a predetermined concentration and / or length, after determining the area size below the peak representing the concentration of cfDNA from mononuclear bodies, this area size can be calibrated by measuring the area size across a reference size range (e.g., any size range described herein) of the control samples and / or molecular weight standards. For example, calibration can be performed to determine the concentration of cfDNA from mononuclear bodies by comparing it to the area size of one or more peaks of the control samples and / or molecular weight standards, including the peak representing the at least one molecular weight standard spot having a predetermined DNA concentration and / or length.

[0040] In some embodiments, the methods described herein further include determining the probability that the test sample originated from a cancer patient based on the concentration of cfDNA corresponding to a size range of single nucleosomes, as determined using the methods described herein. Compared to conventional methods that correlate the total area under the entire cfDNA distribution curve or conventional methods analyzed by a Qubit™ fluorometer (Thermo Fisher Scientific, Q33226), the methods described herein (e.g., CE-based methods analyzing control samples and / or molecular weight standards in the same batch) offer numerous benefits, such as the minimal batch-to-batch variability described above. Furthermore, the methods described herein reduce the risk of contamination by longer DNA fragments (e.g., due to hemolysis during storage or transport), thereby reducing the inclusion of false positive results. Figure 5 As shown, the concentration of cfDNA corresponding to the size range of mononuclear bodies, as determined using the method described in this paper, is positively correlated with cancer incidence, with p-values ​​less than 1 x 10⁻⁶. -6 Therefore, the concentration of cfDNA corresponding to the size range of a single nucleosome can be used as a valuable indicator to determine the probability that a test sample originated from a cancer patient, or to predict the likelihood that a subject whose test sample was collected has cancer. In some embodiments, the concentration of cfDNA corresponding to the size range of a single nucleosome in a test sample from a subject can be compared with the concentration from a cohort of healthy individuals.

[0041] In one aspect, the present invention relates to a method for predicting cancer by determining the concentration of cfDNA within a size range corresponding to mononuclear bodies, wherein the cfDNA is isolated (e.g., extracted using any method described herein) from a test sample (e.g., any tumor sample or healthy sample described herein). The method may include the step of separating plasma from the sample, subsequently extracting cfDNA from the plasma, and quantifying the cfDNA concentration within a size range corresponding to mononuclear bodies by capillary electrophoresis.

[0042] In some embodiments, the concentration of cfDNA corresponding to the size range of a single nucleosome, measured in the subject's test sample, is compared to a reference value (e.g., the concentration of cfDNA corresponding to the size range of a single nucleosome in healthy subjects, or the average concentration of cfDNA corresponding to the size range of a single nucleosome in a group of healthy subjects). For example, if the concentration of cfDNA corresponding to the size range of a single nucleosome, measured in the subject's test sample, is higher than the reference value (e.g., at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least twice as high), the subject is likely to have cancer. In some embodiments, ROC curves can be plotted based on cfDNA concentrations corresponding to the size range of the mononucleosome, and the AUC values ​​can be at least or about 0.65, at least or about 0.66, at least or about 0.67, at least or about 0.68, at least or about 0.69, at least or about 0.70, at least or about 0.71, at least or about 0.72, at least or about 0.73, at least or about 0.74, at least or about 0.75, at least or about 0.76, at least or about 0.77, at least or about 0.78, at least or about 0.79, at least or about 0.80.

[0043] Determination of the proportion of short fragments In one aspect, the present invention relates to a method for determining the proportion of short cfDNA fragments derived from mononuclear bodies in a test sample. In some embodiments, the method described herein includes: (a) isolating cfDNA from the test sample; (b) analyzing the cfDNA by capillary electrophoresis to generate a distribution curve, wherein a control sample and / or molecular weight standard are analyzed in the same batch as the cfDNA; and (c) determining the proportion of short cfDNA fragments derived from mononuclear bodies using the control sample and / or the molecular weight standard as a reference. In some embodiments, the control sample comprises a predetermined DNA concentration. In some embodiments, the molecular weight standard comprises one or more molecular weight standard spots (e.g., DNA molecular weight standard spots), and at least one of the one or more molecular weight standard spots has a predetermined DNA length.

[0044] In some embodiments, the test sample described herein is blood, urine, aqueous humor, or any similar sample from a subject (e.g., a human subject).

[0045] In some embodiments, results obtained from different test samples, control samples, and / or molecular weight standards in the same batch (e.g., distribution curves of cfDNA and such) Figure 2(One or more peaks of the molecular weight standard shown at different time migrations are superimposed for analysis.) Any control samples and / or molecular weight standards described herein can be used to determine the proportion of short cfDNA fragments from mononuclear bodies.

[0046] In some embodiments, the control samples described herein include circulating tumor DNA reference samples prepared by inducing tumor cell apoptosis. Details regarding the preparation of such reference samples can be found, for example, in U.S. Patent Application Publication No. 20230079748A1, which is incorporated herein by reference. In some embodiments, the control samples and / or molecular weight standards described herein include a combination of in vitro synthesized molecular weight standard points and molecular weight standard points obtained from tumor cells as described above.

[0047] In some embodiments, the molecular weight standards described herein include at least one molecular weight standard spot having a predetermined DNA concentration and / or length. In some embodiments, the at least one molecular weight standard spot has a predetermined length derived from an estimated range (e.g., from about 30 bp to about 250 bp) of a peak representing a mononuclear body. In some embodiments, the at least one molecular weight standard spot has a predetermined length of about 30 bp to about 250 bp (e.g., about 30 bp, about 40 bp, about 50 bp, about 60 bp, about 70 bp, about 80 bp, about 90 bp, about 100 bp, about 110 bp, about 120 bp, about 130 bp, about 140 bp, about 150 bp, about 160 bp, about 170 bp, about 180 bp, about 190 bp, about 200 bp, about 210 bp, about 220 bp, about 230 bp, about 240 bp, or about 250 bp). In some embodiments, the length of the at least one molecular weight standard point is approximately 30 bp to approximately 250 bp, approximately 30 bp to approximately 200 bp, approximately 30 bp to approximately 150 bp, approximately 30 bp to approximately 100 bp, approximately 30 bp to approximately 90 bp, approximately 30 bp to approximately 80 bp, approximately 30 bp to approximately 70 bp, approximately 30 bp to approximately 60 bp, approximately 30 bp to approximately 50 bp, approximately 30 bp to approximately 40 bp, approximately 40 bp to approximately 250 bp, approximately 40 bp to approximately 200 bp, approximately 40 bp to approximately 150 bp, approximately 40 bp to approximately 100 bp, approximately 40 bp to approximately 90 bp, approximately 40 bp to approximately 80 bp, approximately 40 bp to approximately 70 bp, approximately 40 bp to approximately 60 bp, approximately 40 bp to approximately 50 bp, approximately 50 bp to approximately 250 bp, approximately 50 bp to approximately 200 bp, approximately 50 bp to approximately 150 bp, approximately 50 bp to approximately 100 bp, approximately 50 bp to approximately 90 bp. 0bp, approximately 50bp to approximately 80bp, approximately 50bp to approximately 70bp, approximately 50bp to approximately 60bp, approximately 60bp to approximately 250bp, approximately 60bp to approximately 200bp, approximately 60bp to approximately 150bp, approximately 60bp to approximately 100bp, approximately 60bp to approximately 90bp, approximately 60bp to approximately 80bp, approximately 60bp to approximately 70bp, approximately 70bp to approximately 250bp, approximately 70bp to approximately 200bp, approximately 70bp to Approximately 150 bp, approximately 70 bp to approximately 100 bp, approximately 70 bp to approximately 90 bp, approximately 70 bp to approximately 80 bp, approximately 80 bp to approximately 250 bp, approximately 80 bp to approximately 200 bp, approximately 80 bp to approximately 150 bp, approximately 80 bp to approximately 100 bp, approximately 80 bp to approximately 90 bp, approximately 90 bp to approximately 250 bp, approximately 90 bp to approximately 200 bp, approximately 90 bp to approximately 150 bp, or approximately 90 bp to approximately 100 bp. In some embodiments, the length of the at least one molecular weight standard point is approximately 100 bp to approximately 250 bp, approximately 100 bp to approximately 240 bp.Approximately 100bp to approximately 230bp, approximately 100bp to approximately 220bp, approximately 100bp to approximately 210bp, approximately 100bp to approximately 200bp, approximately 100bp to approximately 190bp, approximately 100bp to approximately 180bp, approximately 100bp to approximately 170bp, approximately 100bp to approximately 160bp, approximately 100bp to approximately 150bp, approximately 100bp to approximately 140bp, approximately 100bp to approximately 130bp, approximately 100bp to approximately 120bp, approximately 100bp to approximately 110bp, approximately 110bp to approximately 250bp, approximately 110bp to approximately 240bp, approximately 110bp to approximately 230bp, approximately 110bp to approximately 220bp, approximately 110bp to approximately 21 0bp, approximately 110bp to approximately 200bp, approximately 110bp to approximately 190bp, approximately 110bp to approximately 180bp, approximately 110bp to approximately 170bp, approximately 110bp to approximately 160bp, approximately 110bp to approximately 150bp, approximately 110bp to approximately 140bp, approximately 110bp to approximately 130bp, approximately 110bp to approximately 120bp, approximately 120bp to approximately 250bp, approximately 120bp to approximately 240bp, approximately 120bp to approximately 230bp, approximately 120bp to approximately 220bp, approximately 120bp to approximately 210bp, approximately 120bp to approximately 200bp, approximately 120bp to approximately 190bp, approximately 120bp to approximately 180bp, approximately 120bp to Approximately 170bp, approximately 120bp to approximately 160bp, approximately 120bp to approximately 150bp, approximately 120bp to approximately 140bp, approximately 120bp to approximately 130bp, approximately 130bp to approximately 250bp, approximately 130bp to approximately 240bp, approximately 130bp to approximately 230bp, approximately 130bp to approximately 220bp, approximately 130bp to approximately 210bp, approximately 130bp to approximately 200bp, approximately 130bp to approximately 190bp, approximately 130bp to approximately 180bp, approximately 130bp to approximately 170bp, approximately 130bp to approximately 160bp, approximately 130bp to approximately 150bp, approximately 130bp to approximately 140bp, approximately 140bp to approximately 250bp, approximately 14 0bp to approximately 240bp, approximately 140bp to approximately 230bp, approximately 140bp to approximately 220bp, approximately 140bp to approximately 210bp, approximately 140bp to approximately 200bp, approximately 140bp to approximately 190bp, approximately 140bp to approximately 180bp, approximately 140bp to approximately 170bp, approximately 140bp to approximately 160bp, approximately 140bp to approximately 150bp, approximately 150bp to approximately 250bp, approximately 150bp to approximately 240bp, approximately 150bp to approximately 230bp, approximately 150bp to approximately 220bp, approximately 150bp to approximately 210bp, approximately 150bp to approximately 200bp, approximately 150bp to approximately 190bp, approximately 150bp to approximately 180bp.Approximately 150bp to approximately 170bp, approximately 150bp to approximately 160bp, approximately 160bp to approximately 250bp, approximately 160bp to approximately 240bp, approximately 160bp to approximately 230bp, approximately 160bp to approximately 220bp, approximately 160bp to approximately 210bp, approximately 160bp to approximately 200bp, approximately 160bp to approximately 190bp, approximately 160bp to approximately 180bp, approximately 160bp to approximately 170bp, approximately 170bp to approximately 250bp. 0bp, approximately 170bp to approximately 240bp, approximately 170bp to approximately 230bp, approximately 170bp to approximately 220bp, approximately 170bp to approximately 210bp, approximately 170bp to approximately 200bp, approximately 170bp to approximately 190bp, approximately 170bp to approximately 180bp, approximately 180bp to approximately 250bp, approximately 180bp to approximately 240bp, approximately 180bp to approximately 230bp, approximately 180bp to approximately 220bp, approximately 180bp to Approximately 210bp, approximately 180bp to approximately 200bp, approximately 180bp to approximately 190bp, approximately 190bp to approximately 250bp, approximately 190bp to approximately 240bp, approximately 190bp to approximately 230bp, approximately 190bp to approximately 220bp, approximately 190bp to approximately 210bp, approximately 190bp to approximately 200bp, approximately 200bp to approximately 250bp, approximately 200bp to approximately 240bp, approximately 200bp to approximately 230bp, approximately 200 From approximately 220bp to approximately 200bp to approximately 210bp, from approximately 210bp to approximately 250bp, from approximately 210bp to approximately 240bp, from approximately 210bp to approximately 230bp, from approximately 210bp to approximately 220bp, from approximately 220bp to approximately 250bp, from approximately 220bp to approximately 240bp, from approximately 220bp to approximately 230bp, from approximately 230bp to approximately 250bp, from approximately 230bp to approximately 240bp, or from approximately 240bp to approximately 250bp.

[0048] Based on NGS-based analysis results ( Figure 1 A) The initiation position of a mononuclear body is typically around 30-100 bp, and the termination position is typically around 250 bp. Therefore, by overlaying the results obtained from capillary electrophoresis plots of the test sample and control samples and / or molecular weight standards, a person skilled in the art can identify a peak in the cfDNA distribution curve that migrates at substantially the same time as the at least one molecular weight standard point, and that this peak represents a migrating mononuclear body. Therefore, the methods described herein may include identifying such a peak by recognizing its initiation and / or termination position based on the distribution curve.

[0049] In some embodiments, the molecular weight standards described herein also include at least two molecular weight standard spots (e.g., DNA molecular weight standard spots) having predetermined concentrations and / or lengths. In some embodiments, the lengths of these two molecular weight standard spots are designed such that the termination position of the peak representing a mononuclear body lies between them. Similarly, the lengths of these two molecular weight standard spots can be designed such that the start position of the peak representing a mononuclear body lies between them. In some embodiments, the lengths of these two molecular weight standard spots are designed such that either or both of the start and termination positions of the peak representing a mononuclear body lie between them. By superimposing the results obtained from capillary electrophoresis plots of the test sample and the molecular weight standards, those skilled in the art can identify the start and / or termination positions of the peak representing a migrating mononuclear body by identifying the lowest point (e.g., a local minimum) of the cfDNA distribution curve between the first and second molecular weight standard spots (e.g., any of the first and second molecular weight standard spots described herein). Therefore, the methods described herein may include identifying the start and / or termination positions of the aforementioned peaks based on the distribution curve.

[0050] The method described in this article can be further optimized to determine the concentration of cfDNA from mononuclear bodies more accurately. For example... Figure 4 As shown, the cfDNA distribution curve may be downward shifted, leading to an underestimation of the cfDNA concentration from mononuclear bodies. Similarly, the cfDNA distribution curve may be upward shifted, leading to an overestimation of the cfDNA concentration from mononuclear bodies. To overcome these problems, the method described herein may further include identifying the start and / or end positions of the peak representing the aforementioned migration of mononuclear bodies, and can more accurately measure the area under that peak. For example, the cfDNA concentration from mononuclear bodies can be represented by the size of a first area defined by the distribution curve and a straight line connecting the start and end positions. Similarly, the cfDNA concentration shorter than the at least one molecular weight standard point (e.g., a molecular weight standard point with a predetermined length of 150 bp) can be represented by the size of a second area defined by the distribution curve, the predetermined molecular weight standard point representing the at least one molecular weight standard point, and a straight line connecting the start and end positions. The proportion of short cfDNA fragments from mononuclear bodies described herein can be calculated by dividing the size of the second area by the size of the first area.

[0051] As another example, the concentration of cfDNA from mononuclear bodies can be represented by the size of a first area defined by a cfDNA distribution curve within a size range (e.g., any size range described herein) and a straight line connecting two adjacent local minima of that distribution curve. In some embodiments, the concentration of short fragments from mononuclear bodies can be represented by a second area defined by a distribution curve, a predetermined molecular weight standard point representing the upper limit of the short fragment size range (e.g., 150 bp), and a straight line connecting two adjacent local minima of the cfDNA distribution curve within said size range. Because the control samples and / or molecular weight standards described herein may include at least one molecular weight standard point having a predetermined concentration and / or length, after determining the aforementioned area sizes, they can be calibrated by measuring the area size across a reference size range (e.g., any size range described herein) of the control samples and / or molecular weight standards. For example, calibration can be performed by comparing the area size with one or more peaks of the control samples and / or molecular weight standards (including a peak representing said at least one molecular weight standard point having a predetermined DNA concentration and / or length) to determine the cfDNA concentration from mononuclear bodies. The proportion of short cfDNA fragments from mononuclear bodies described in this article can be calculated by dividing the size of the second area by the size of the first area.

[0052] In some embodiments, the methods described herein further include determining the probability that a test sample originated from a cancer patient based on the proportion of short cfDNA fragments from mononuclear bodies determined using the methods described herein. The methods described herein (e.g., CE-based methods for analyzing control samples and / or molecular weight standards in the same batch) achieve minimal systematic error when comparing the determined proportions of short cfDNA fragments from mononuclear bodies (e.g., between subjects with cancer and healthy subjects).

[0053] like Figure 7 As shown in Figure A, the proportion of short cfDNA fragments from mononuclear bodies, determined using the method described in this paper, is positively correlated with cancer incidence, with a p-value less than 1 x 10⁻⁶. -5 Therefore, the proportion of short cfDNA fragments from mononuclei can be used as a valuable indicator to determine the probability that a test sample originated from a cancer patient, or to predict the likelihood that a subject whose test sample was collected has cancer. In some embodiments, the proportion of short cfDNA fragments from mononuclei in a subject's test sample can be compared to the proportion from a cohort of healthy individuals.

[0054] In one aspect, the present invention relates to a method for predicting cancer by determining the proportion of short cfDNA fragments from mononuclear bodies, wherein the cfDNA is isolated (e.g., extracted using any of the methods described herein) from a test sample (e.g., any tumor sample or healthy sample described herein). The method may include the step of separating plasma from the sample, subsequently extracting cfDNA from the plasma, and calculating the proportion of short cfDNA fragments from mononuclear bodies using capillary electrophoresis.

[0055] In some embodiments, the proportion of cfDNA fragments from mononuclear bodies measured in a subject's test sample is compared to a reference value (e.g., the proportion of cfDNA fragments from mononuclear bodies in healthy subjects, or the average proportion of cfDNA fragments from mononuclear bodies in a group of healthy subjects). For example, if the proportion of cfDNA fragments from mononuclear bodies measured in a subject's test sample is higher than a reference value (e.g., at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least twice as high), the subject is likely to have cancer. In some embodiments, the proportion of cfDNA fragments from mononuclear bodies measured in a subject's (e.g., a subject with cancer) test sample is at least 10%, at least 11%, at least 12%, at least 13%, at least 14%, at least 15%, at least 16%, at least 17%, at least 18%, at least 19%, or at least 20%. In some embodiments, the proportion of short cfDNA fragments from mononuclear bodies measured in the test sample of a subject (e.g., a healthy subject) is less than 20%, less than 19%, less than 18%, less than 17%, less than 16%, less than 15%, less than 14%, less than 13%, less than 12%, less than 11%, or less than 10%. In some embodiments, an ROC curve can be plotted based on the proportion of short cfDNA fragments from mononuclear bodies, and the AUC value can be at least or about 0.65, at least or about 0.66, at least or about 0.67, at least or about 0.68, at least or about 0.69, at least or about 0.70, at least or about 0.71, at least or about 0.72, at least or about 0.73, at least or about 0.74, at least or about 0.75, at least or about 0.76, at least or about 0.77, at least or about 0.78, at least or about 0.79, or at least or about 0.80.

[0056] Measurement of markers in samples For decades, large-scale blood-based immunological measurements of protein biomarkers (PTMs) have been performed clinically in seemingly healthy individuals for cancer screening; for example, alpha-fetoprotein (AFP) for liver cancer, CA125 for ovarian cancer, CA15-3 for breast cancer, CA19-9 for pancreatic cancer, CA72-4 for ovarian cancer, carcinoembryonic antigen (CEA) for gastrointestinal cancers, and CYFRA21-1 for breast cancer. These methods offer significant advantages over many other clinical diagnostic methods (endoscopy, imaging, etc.), including their non-invasiveness, automation, and relatively low cost. However, their low sensitivity in early cancer detection limits their widespread use for screening purposes in the general population.

[0057] Previous studies have shown that PTM combinations are superior to single biomarkers in the early detection of colorectal cancer, lung cancer, breast cancer, liver cancer, gastric cancer, pancreatic cancer, ovarian cancer, and esophageal cancer. Several reports have also demonstrated that combined PTM combinations can be used to detect several cancer types simultaneously. However, different cancer types often exhibit different serological characteristics. As sample sizes increase, test results may also become more complex, and traditional statistical methods may be unable to handle such large datasets. Furthermore, traditional clinical methods that simultaneously detect multiple PTMs and use a single threshold to evaluate results can lead to a cumulative false positive rate and unnecessary clinical diagnostic testing. Therefore, they are not suitable for screening large populations with asymptomatic cases. AI is a good analytical approach to solve classification challenges by identifying latent patterns from complex data. Over the past decade, AI technology has made significant contributions to this advanced field, playing a crucial role in medical and healthcare research. AI is considered a valuable tool for transforming the future of healthcare and precision oncology. Several novel algorithms have shown promising results in the accurate detection and characterization of suspected lesions.

[0058] Details on measuring markers (e.g., biomarkers) in a sample can be found, for example, Chinese Patent Application Publication No. 117831690A, which is incorporated herein by reference in its entirety.

[0059] biomarkers Before measurement, a set of biomarkers needs to be selected for the specific cancer being screened. Many biomarkers for diseases, including cancer, are known, and known combinations can be selected, or as described in the examples. The combination can be selected based on measurements of individual biomarkers in a retrospective clinical sample, wherein the combination is generated based on empirical data for the desired disease, such as cancer, and preferably pancreatic cancer, ovarian cancer, liver cancer, lung cancer, stomach cancer, colorectal cancer, lymphoma, esophageal cancer, or breast cancer.

[0060] Examples of usable biomarkers include molecules detectable in bodily fluid samples, such as antibodies, antigens, small molecules, proteins, hormones, enzymes, genes, etc. However, the use of tumor antigens has many advantages because they have been widely used for many years, and many have validated and standardized test kits available for use with the aforementioned automated immunoassay platforms.

[0061] In certain embodiments, a set of biomarkers is selected based on their association with specific cancer types. For example, AFP is a specific biomarker for liver cancer. Additionally, alpha-fetoprotein (AFP) can be used as a biomarker for liver cancer (e.g., hepatocellular carcinoma), CA125 for ovarian cancer, CA15-3 for breast cancer, CA19-9 for pancreatic cancer, CA72-4 for ovarian cancer, carcinoembryonic antigen (CEA) for gastrointestinal cancers, and CYFRA21-1 for breast cancer.

[0062] In some embodiments, the set of markers may include markers associated with a selection of cancers including pancreatic cancer, ovarian cancer, liver cancer, lung cancer, stomach cancer, colorectal cancer, lymphoma, esophageal cancer, prostate cancer, or breast cancer.

[0063] As a design option, the combination can include any number of biomarkers in order to seek, for example, to maximize the specificity or sensitivity of the assay. Thus, as a design option, the assay of interest may require the presence of at least one of two, three, four, five, six, seven, eight, nine, ten or more biomarkers.

[0064] Therefore, in one embodiment, the biomarker combination may comprise at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten or more different biomarkers. In one embodiment, the biomarker combination comprises about two to ten different biomarkers. In another embodiment, the biomarker combination comprises about four to eight different biomarkers. In yet another embodiment, the biomarker combination comprises about seven different biomarkers. In still another embodiment, the biomarker combination comprises about ten different biomarkers. Typically, the sample is used for assays, and the results can be a series of numbers reflecting the presence and level (e.g., concentration, quantity, activity, etc.) of each biomarker in the combination in the sample.

[0065] In some embodiments, the biomarker combination described herein includes one or more biomarkers selected from AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA21-1, ProGRP, SCCA, and PSA. In some embodiments, the biomarker combination described herein includes at least seven different biomarkers selected from AFP, CA125, CA15-3, CA19-9, CEA, CYFRA21-1, ProGRP, SCCA, and PSA. In some embodiments, the biomarker combination described herein includes AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, and CYFRA21-1. In some embodiments, the subject is male, and the biomarker combination described herein includes AFP, CA125, CA15-3, CA19-9, CEA, CYFRA21-1, ProGRP, SCCA, and PSA; or in some embodiments, the subject is female, and the biomarker combination described herein includes AFP, CA125, CA15-3, CA19-9, CEA, CYFRA21-1, ProGRP, and SCCA. In some embodiments, the biomarker combination described herein includes CEA, CYFRA21-1, SCCA, and ProGRP, and the cancer is lung cancer. In some embodiments, the biomarker combination described herein includes AFP, CA125, CA15-3, CA19-9, CEA, and CYFRA21-1. In some embodiments, the subject is male, and the combination of biomarkers described herein includes AFP, CA125, CA15-3, CA19-9, CEA, CYFRA21-1, and PSA; or in some embodiments, the subject is female, and the combination of biomarkers described herein includes AFP, CA125, CA15-3, CA19-9, CEA, and CYFRA21-1. In some embodiments, the subject is male, and the combination of biomarkers described herein includes AFP, CA125, CA15-3, CA19-9, CEA, CYFRA21-1, SCCA, and PSA; or in some embodiments, the subject is female, and the combination of biomarkers described herein includes AFP, CA125, CA15-3, CA19-9, CEA, CYFRA21-1, and SCCA. Details of the biomarkers can be found, for example, in Chinese Patent Application Publication No. 117831690A, which is incorporated herein by reference in its entirety.

[0066] Multi-cancer early detection (MCED) model Embodiments of this invention provide techniques for developing machine learning systems for multi-cancer early detection (MCED) testing, enabling the identification of more than one type of cancer from a single test sample (e.g., a single blood sample) with high sensitivity and accuracy. The machine learning system can be trained using one or more machine learning algorithms to distinguish between individuals with and without cancer. This invention is partly based on observational studies (e.g., Figure 8 As shown in the figure, the two cfDNA characteristics (including the concentration and proportion of short fragments of cfDNA from mononuclei) measured using the methods described herein can complement each other in cancer detection.

[0067] In one aspect, the present invention provides a computer-implemented method for early detection of the presence of cancer in a patient, the computer-implemented method comprising: (a) determining the concentration of cfDNA in a test sample of a subject within a size range corresponding to a single nucleosome; (b) determining the proportion of short cfDNA fragments from single nucleosomes in the test sample of the subject; (c) selecting multiple parameters to input into a machine learning system, wherein the multiple parameters include cfDNA concentration and the proportion of short cfDNA fragments from single nucleosomes; (d) training the machine learning system using a machine learning algorithm selected from random forest (RF), generalized linear model (GLM), support vector machine (SVM), and / or gradient boosting machine (GBM); and (e) determining a cancer prediction score, wherein a higher cancer prediction score indicates a higher probability that the subject has cancer.

[0068] In one aspect, the present invention provides a computer-implemented method for early detection of the presence of cancer in a patient, the computer-implemented method comprising: (a) quantifying the levels of a set of biomarkers in a subject's test sample (e.g., a blood sample), wherein the set of biomarkers includes one or more biomarkers selected from AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA21-1, SCCA, ProGRP, and PSA; (b) determining the concentration of cfDNA in the subject's test sample corresponding to a single nucleosome size range; and (c) determining the level of the subject's test sample. (d) The proportion of cfDNA fragments from mononuclear bodies in the test sample; (e) Selecting multiple parameters to input into the machine learning system, wherein the multiple parameters include the level of one or more biomarkers, the concentration of cfDNA corresponding to the size range of mononuclear bodies, and the proportion of cfDNA fragments from mononuclear bodies; (f) Training the machine learning system using a machine learning algorithm selected from random forest (RF), generalized linear model (GLM), support vector machine (SVM) and / or gradient boosting machine (GBM); and (f) Determining a cancer prediction score, wherein a higher cancer prediction score indicates a higher probability that the subject has cancer.

[0069] In some embodiments, a generalized linear model (GLM) is used. GLM is a flexible generalization of ordinary linear regression. GLM generalizes linear regression by allowing the linear model to be associated with the response variable via a link function and by allowing the variance of each measurement to be a function of its predicted value. GLM is formulated as a method to unify various other statistical models, including linear regression, logistic regression, and Poisson regression. In some embodiments, the iteratively reweighted least squares method is used to perform maximum likelihood estimation (MLE) of the model parameters. MLE remains popular and is the default method on many statistical computing packages. Other methods, including Bayesian regression and least squares fitting for variance-stabilized responses, have also been developed.

[0070] In some embodiments, the methods described herein involve gradient boosting. Gradient boosting is a machine learning technique used in fields such as regression and classification. It provides a predictive model as an ensemble of weak predictive models, typically decision trees. In some embodiments, the gradient-boostedtrees model is constructed in a stage-wise manner, like other boosting methods, but it generalizes to other methods by allowing optimization of arbitrarily differentiable loss functions.

[0071] In some embodiments, the methods described herein involve random forests (RF). Random forests are an ensemble learning method for classification, regression, and other tasks that operates by building a large number of decision trees during training. For classification tasks, the output of a random forest is the class selected by the majority of trees. For regression tasks, it returns the average or average prediction of the individual trees. Random decision forests correct the tendency of decision trees to overfit their training set.

[0072] In some embodiments, the methods described herein involve Support Vector Machines (SVMs, also known as Support Vector Networks). SVMs are supervised learning models with associated learning algorithms that analyze data for classification and regression analysis. Given a set of training examples, each labeled as belonging to one of two classes, the SVM training algorithm builds a model that assigns new examples to one class or the other, making it a non-probabilistic binary linear classifier (although methods such as Platt scaling exist to use SVMs in probabilistic classification environments). The SVM maps training examples to points in a space to maximize the gap width between the two classes. New examples are then mapped to the same space and their class is predicted based on which side of the gap they fall on. In addition to performing linear classification, SVMs can also efficiently perform non-linear classification using the so-called kernel trick, implicitly mapping their input to a high-dimensional feature space.

[0073] In some embodiments, the methods described herein include one or more ensemble models that improve overall prediction performance by combining models trained using different algorithms, parameters, and / or training data. Ensemble models do not rely on a single model but rather leverage the diversity of multiple models to improve prediction accuracy and robustness.

[0074] Different cancer types require specific panels of PTMs for cancer screening. However, traditional clinical methods rely solely on a single threshold for each PTM, which presents challenges when combining results from multiple PTMs, leading to a build-up of false positives with increasing biomarkers. Nevertheless, in some embodiments, machine learning systems are trained with one or more machine learning algorithms to process data from multiple cancer types together, thereby distinguishing cancerous individuals from non-cancer individuals by calculating a Probability of Cancer (POC) index based on two cfDNA features described herein (e.g., concentration and proportion of short fragments of cfDNA from mononuclei); expression of multiple PTMs; and / or clinical baseline information including individual sex and age. By integrating multiple test results into a single outcome, it significantly reduces the false positive rate while maintaining the combined sensitivity of multiple tests and eliminating differences in PTM levels between different demographic groups (e.g., various age groups and sexes). Therefore, this method offers good performance for the simultaneous detection of multiple cancer types. After determining whether a subject is likely to have cancer using the machine learning system, the system can be used to treat the subject's cancer, monitor disease progression, determine the effectiveness of treatment, and adjust treatment strategies. This computer-implemented method enables earlier cancer detection and more timely assessment of treatment effectiveness, thereby reducing the likelihood of missing the optimal treatment window and lowering cancer mortality rates, ultimately improving the quality of life for cancer patients. The computer-implemented method described in this paper is powered by AI technology to significantly reduce the false positive rate.

[0075] The method described in this paper significantly outperforms traditional clinical approaches, representing a novel blood-based test for early detection of multiple cancer types (MCED) that is non-invasive, simple, efficient, and robust. Furthermore, the method described is affordable and readily available, requiring only a blood draw at screening sites, making it acceptable and sustainable in low- and middle-income countries (LMICs). This invention provides a protein detection method that integrates measurements of a selected set of protein markers (e.g., seven or ten) with individual clinical information and is greatly enhanced by AI technology, making it more practical in LMICs.

[0076] In some embodiments, the training dataset includes at least 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10000 subjects. In some embodiments, true positive cancer patients account for at least 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, or 80% of the sample size. In some embodiments, the training dataset has no more than 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10000 subjects.

[0077] In some embodiments, the methods described herein can achieve a sensitivity of at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%. In some embodiments, the sensitivity is about 50% to about 100%, about 55% to about 100%, about 60% to about 100%, about 65% to about 100%, about 70% to about 100%, about 75% to about 100%, about 80% to about 100%, about 85% to about 100%, about 90% to about 100%, about 95% to about 100%, about 50% to about 95%, about 55% to about 95%, about 60% to about 95%, about 65% to about 95%, about 70% to about 95%, about 75% to about 95%, about 80% to about 95%, about 85% to about 95%, about 90% to about 95%, about 50% to about 90%, about 55% to about 90%, about 60% to about 90%, about 65% to about 90%, about 70% to about 90%, about 75% to about 90%, about 80% to about 90%, about 85%. From approximately 90%, approximately 50% to approximately 85%, approximately 55% to approximately 85%, approximately 60% to approximately 85%, approximately 65% ​​to approximately 85%, approximately 70% to approximately 85%, approximately 75% to approximately 85%, approximately 80% to approximately 85%, approximately 50% to approximately 80%, approximately 55% to approximately 80%, approximately 60% to approximately 80%, approximately 65% ​​to approximately 80%, approximately 70% to approximately 80%, approximately 75% to approximately 80%, approximately 50% to approximately 75%, approximately 55% to approximately 75%, approximately 60% to approximately 75%, approximately 65% ​​to approximately 75%, approximately 70% to approximately 75%, approximately 50% to approximately 70%, approximately 55% to approximately 70%, approximately 60% to approximately 70%, approximately 65% ​​to approximately 70%, approximately 50% to approximately 65%, approximately 55% to approximately 65%, approximately 60% to approximately 65%, approximately 50% to approximately 60%, approximately 55% to approximately 60%, or approximately 50% to approximately 55%.

[0078] In some embodiments, the methods described herein can achieve at least 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% specificity. In some embodiments, specificity is about 75% to about 100%, about 80% to about 100%, about 85% to about 100%, about 90% to about 100%, about 95% to about 100%, about 75% to about 95%, about 80% to about 95%, about 85% to about 95%, about 90% to about 95%, about 75% to about 90%, about 80% to about 90%, about 85% to about 90%, about 75% to about 85%, about 80% to about 85%, or about 75% to about 80%.

[0079] In some embodiments, the methods described herein involve a cross-validation process. In some embodiments, the cross-validation process is a 50%, 60%, 70%, 80%, 90%, 100%, 15%, 20%, 25%, 30%, 35%, or 40% cross-validation process. In some embodiments, the cross-validation process is repeated at least 10, 20, 30, 40, 50, or 100 times. In some embodiments, compared to conventional methods for detecting cancer (e.g., methods based on predetermined reference ranges for each biomarker), the methods described herein can reduce the false positive rate by at least 10%, 20%, 30%, 40%, or 50%.

[0080] Furthermore, in some cases, certain protein biomarkers associated with multiple cancer types have a greater contribution or higher weight in the model. Conversely, certain biomarkers that are highly specific to a particular cancer type (e.g., AFP, which is specific for liver cancer detection) contribute relatively less. This leads to situations where, in some cases, a highly specific protein biomarker for a particular cancer type exhibits abnormally high levels (while other protein biomarkers remain normal), and the MCED model may predict a lower cancer probability (POC) index.

[0081] To address this issue, an outlier analysis method was developed to predict patients with these types of cancer. Outlier analysis focuses on identifying and analyzing cases where highly specific cancer biomarkers show aberrant expression levels compared to normal cases. By incorporating this method into the MCED model, the identification of cancer patients who may exhibit unique biomarker expression was improved, providing more accurate predictions and insights for early diagnosis of multiple cancer types. Here, we determine the cutoff value for outlier analysis using three methods based on over 6000 non-cancer samples.

[0082] 1) Boxplot method: The boxplot method is used to identify outliers by plotting the expression of protein biomarkers from normal control samples. The boxplot shows the quartile range of the data, and observations exceeding the upper quartile plus 1.5 times the interquartile range can be considered as cutoff values ​​for outliers.

[0083] 2) Modified Z-Score: Because certain non-cancerous diseases can also lead to elevated protein biomarker expression levels in the normal control cohort, protein expression levels in the normal control cohort exhibit skewness. Therefore, the modified Z-score accounts for data skewness by dividing the difference between the observed value and the median by the median absolute difference (MAD). Protein biomarker expression with a modified Z-score > 10 is defined as the cutoff value for outliers.

[0084] 3) Percentile: The percentile method compares an observation to the percentile of the data. Observations that exceed the 99th percentile of the normal cohort can be considered as outliers.

[0085] Based on the cutoff values ​​obtained using the three methods described above, the maximum value is selected as the final high-anomaly cutoff value. If the expression level of a specific biomarker in a test sample is greater than the corresponding anomaly cutoff value, the sample can be predicted to be a cancer patient. By developing this outlier analysis method, we aim to enhance the identification of cancer patients with significantly abnormal levels of a particular cancer-specific protein biomarker. By effectively predicting these special cases, we can provide valuable insights into the potential presence of specific types of cancer and assist in early detection.

[0086] Therefore, in some embodiments, the methods described herein further include outlier analysis. In some embodiments, the outlier analysis described herein involves determining a cutoff value based on at least 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 20,000, 30,000, 40,000, or 50,000 non-cancer samples. In some embodiments, the cutoff value is determined by selecting the maximum value obtained from the box plot method, corrected Z-score, and / or percentiles described herein. In some embodiments, a cutoff value may be determined for each biomarker described herein. In some embodiments, outlier analysis includes comparing the quantitative level (e.g., expression level) of a selected biomarker (e.g., any biomarker described herein) with its corresponding cutoff value determined herein. For example, if the quantitative level of a selected biomarker is higher than the corresponding cutoff value, the patient has a high probability of having cancer. In some embodiments, outlier analysis is performed by the following steps: (a) determining a cutoff value for each biomarker (e.g., by box plot, adjusted Z-score, and / or percentile), and (b) comparing the quantitative level of each biomarker with its corresponding cutoff value. In some embodiments, a higher quantitative level of a biomarker relative to its corresponding cutoff value indicates a higher probability that the subject has cancer.

[0087] Treatment The computer-implemented method described herein may further include additional steps: after determining whether a subject is likely to have cancer using the machine learning system implemented herein, treating the subject's cancer, assessing the effectiveness of the treatment more promptly, reducing the likelihood of missing the optimal treatment window, reducing the rate of tumor volume increase over time in the subject, reducing the risk of metastasis, and / or reducing the risk of additional metastasis in the subject. In some embodiments, the treatment may stop, slow, delay, or inhibit cancer progression. In some embodiments, the treatment may result in a reduction in the number, severity, and / or duration of one or more symptoms of cancer in the subject. In some embodiments, the compositions and methods disclosed herein may be used to treat patients at risk of developing cancer.

[0088] These treatments can typically include, for example, surgery, chemotherapy, radiation therapy, hormone therapy, targeted therapy, and / or combinations thereof. The choice of treatment depends on the type, location, and grade of the cancer, as well as the patient's health condition and preferences. In some embodiments, the treatment is chemotherapy or chemoradiotherapy.

[0089] In some embodiments, the present invention relates to a method for determining whether a postoperative patient should receive treatment. In some embodiments, the method includes determining a cancer prediction score, wherein a high cancer prediction score (e.g., greater than 0.8, 0.81, 0.82, 0.83, 0.84, 0.85, 0.86, 0.87, 0.88, 0.89, 0.9, 0.91, 0.92, 0.93, 0.94, or 0.95) indicates that the patient should receive treatment postoperatively, while a low cancer prediction score (e.g., not greater than 0.8, 0.81, 0.82, 0.83, 0.84, 0.85, 0.86, 0.87, 0.88, 0.89, 0.9, 0.91, 0.92, 0.93, 0.94, or 0.95) indicates that the patient does not require treatment.

[0090] In one aspect, the invention is characterized by a method of administering a therapeutically effective amount of a therapeutic agent to a subject in need (e.g., a subject who has, or has been identified as having, or has been diagnosed with, cancer). In some embodiments, the subject has, for example, breast cancer (e.g., triple-negative breast cancer), carcinoid tumor, cervical cancer, endometrial cancer, glioma, head and neck cancer, liver cancer, lung cancer, small cell lung cancer, lymphoma, melanoma, ovarian cancer, pancreatic cancer, prostate cancer, kidney cancer, colorectal cancer, stomach cancer, testicular cancer, thyroid cancer, bladder cancer, urethral cancer, or a hematologic malignancy. In some embodiments, the cancer is unresectable melanoma or metastatic melanoma, non-small cell lung cancer (NSCLC), small cell lung cancer (SCLC), bladder cancer, or metastatic hormone-resistant prostate cancer. In some embodiments, the subject has a solid tumor. In some embodiments, the cancer is head and neck squamous cell carcinoma (SCCHN), renal cell carcinoma (RCC), triple-negative breast cancer (TNBC), or colorectal cancer. In some embodiments, the subject has triple-negative breast cancer (TNBC), stomach cancer, urothelial carcinoma, Merkel cell carcinoma, or head and neck cancer. In some embodiments, the subject has pancreatic cancer, ovarian cancer, liver cancer, lung cancer, stomach cancer, colorectal cancer, lymphoma, esophageal cancer, or breast cancer.

[0091] As used herein, an "effective amount" means an amount or dose sufficient to produce a beneficial or desired outcome, including stopping, slowing, delaying, or inhibiting the progression of a disease (e.g., cancer). Effective amounts will vary depending on factors such as the age and weight of the subject to be administered the therapeutic agent, the severity of symptoms, and the route of administration, and therefore can be determined on an individual basis. Effective amounts may be administered in one or more administrations. For example, an effective amount is an amount sufficient to improve, stop, stabilize, reverse, inhibit, slow, and / or delay the progression of cancer in a patient, or an amount sufficient to improve, stop, stabilize, reverse, slow, and / or delay the proliferation of cells in vitro (e.g., biopsy cells, any cancer cells or cell lines described herein, e.g., cancer cell lines)).

[0092] In some embodiments, the methods described herein can be used to monitor disease progression, determine the effectiveness of treatment, and adjust treatment strategies. For example, cell-free DNA (cfDNA) can be collected from a subject to detect cancer, and this information can also be used to select appropriate treatment for the subject. Cell-free DNA can be collected from the subject after treatment. Analysis of these cfDNAs can be used to monitor disease progression, determine the effectiveness of treatment, and / or adjust treatment strategies. In some embodiments, these results are then compared with earlier results. In some embodiments, a sharp increase in circulating tumor DNA indicates tumor cell apoptosis, which may suggest that the treatment is effective.

[0093] In some embodiments, the therapeutic agent may comprise one or more inhibitors selected from: B-Raf inhibitors, EGFR inhibitors, MEK inhibitors, ERK inhibitors, K-Ras inhibitors, c-Met inhibitors, anaplastic lymphoma kinase (ALK) inhibitors, phosphatidylinositol 3-kinase (PI3K) inhibitors, Akt inhibitors, mTOR inhibitors, PI3K / mTOR dual inhibitors, Bruton's tyrosine kinase (BTK) inhibitors, and isocitrate dehydrogenase 1 (IDH1) and / or isocitrate dehydrogenase 2 (IDH2) inhibitors. In some embodiments, an additional therapeutic agent is an indoleamine 2,3-dioxygenase-1 (IDO1) inhibitor (e.g., epacadostat). In some embodiments, the therapeutic agent may comprise one or more inhibitors selected from: HER3 inhibitors, LSD1 inhibitors, MDM2 inhibitors, BCL2 inhibitors, CHK1 inhibitors, inhibitors activating the Hedgehog signaling pathway, and agents that selectively degrade estrogen receptors.

[0094] In some embodiments, the therapeutic agent may comprise one or more therapeutic agents selected from: Trabectedin, nab-paclitaxel, Trebananib, Pazopanib, Cediranib, Palbociclib, everolimus, fluoropyrimidine, IFL, regorafenib, Reolysin, Alimta (pemetrexed), Zykadia (ceritinib), Sutent (sunitinib), temsirolimus, axitinib, sorafenib, Votrient (pazopanib). Panib, IMA-901, AGS-003, cabozantinib, vinflunine, Hsp90 inhibitors, Ad-GM-CSF, temozolomide, IL-2, IFNa, vinblastine, thalidomide, dacarbazine, cyclophosphamide, lenalidomide, azacytidine, bortezomid, amrubicin, carfilzomib, pralatrexate, and enzastaurin.

[0095] In some embodiments, the therapeutic agent may comprise one or more therapeutic agents selected from: adjuvants, TLR agonists, tumor necrosis factor (TNF) α, IL-1, HMGB1, IL-10 antagonists, IL-4 antagonists, IL-13 antagonists, IL-17 antagonists, HVEM antagonists, ICOS agonists, therapies targeting CX3CL1, therapies targeting CXCL9, therapies targeting CXCL10, therapies targeting CCL5, LFA-1 agonists, ICAM1 agonists, and selectin agonists.

[0096] In some embodiments, the subject is administered carboplatin, albumin-bound paclitaxel, paclitaxel, cisplatin, pemetrexed, gemcitabine, FOLFOX, or FOLFIRI.

[0097] In some embodiments, the therapeutic agent is an antibody or its antigen-binding fragment. In some embodiments, the therapeutic agent is an antibody that specifically binds to PD-1, CTLA-4, BTLA, PD-L1, CD27, CD28, CD40, CD47, CD137, CD154, TIGIT, TIM-3, GITR, or OX40. In some embodiments, the therapeutic agent is an anti-PD-1 antibody, an anti-OX40 antibody, an anti-PD-L1 antibody, an anti-PD-L2 antibody, an anti-LAG-3 antibody, an anti-TIGIT antibody, an anti-BTLA antibody, an anti-CTLA-4 antibody, or an anti-GITR antibody. In some embodiments, the therapeutic agent is an anti-CTLA4 antibody (e.g., ipilimumab), an anti-CD20 antibody (e.g., rituximab), an anti-EGFR antibody (e.g., cetuximab), an anti-CD319 antibody (e.g., elotuzumab), or an anti-PD1 antibody (e.g., nivolumab).

[0098] Systems, software and interfaces The computer-implemented methods described herein (e.g., quantifying cfDNA concentration, determining cancer risk, determining short fragment proportions, determining cancer prediction scores, etc.) can be implemented by a computer, processor, software, module, or other means. The methods described herein can be computer-implemented, and all or part of the methods may sometimes be executed by one or more processors. Embodiments relating to the methods described herein are applicable to the same or related processes implemented by instructions in the systems, apparatus, and computer program products described herein. In some embodiments, the processes and methods described herein are performed by automated methods. In some embodiments, automated methods are embodied in software, modules, processors, peripheral devices, and / or means including the foregoing for determining sequence reads, counting, mapping, aligning sequence tags, height, mapping, normalization, comparison, range setting, classification, adjustment, plotting, results, conversion, and identification. As used herein, software means computer-readable program instructions that, when executed by a processor, perform the computer operations described herein.

[0099] Information from cfDNA, biomarkers, and maps, along with other relevant information, derived from subjects (e.g., control subjects, patients, or subjects suspected of having tumors), can be analyzed and processed to determine the presence of genetic variations. Sequence reads and counts are sometimes referred to as “data” or “datasets.” In some embodiments, data or datasets can be characterized by one or more features or variables. In some embodiments, a sequencing device is included as part of the system. In some embodiments, the system includes a computing device and a sequencing device, wherein the sequencing device is configured to receive physical nucleic acids and generate sequence reads, while the computing device is configured to process the reads from the sequencing device. The computing device is sometimes configured to determine the presence of genetic variations (e.g., copy number variations, mutations) from the sequence reads. In some embodiments, the system includes a capillary electrophoresis mapping device.

[0100] The implementation of the subject matter and functional operations described herein can be implemented in digital electronic circuits, in tangibly embodied computer software or firmware, in computer hardware (including the structures described herein and their equivalents), or in a combination of one or more of these structures. The implementation of the subject matter described herein can be implemented as one or more computer programs, i.e., one or more computer program instruction modules encoded on a tangible program carrier for execution by a processing device or for controlling the operation of a processing device. Alternatively or additionally, program instructions can be encoded on propagating signals, which are artificially generated signals, such as machine-generated electrical, optical, or electromagnetic signals, generated to encode information for transmission to a suitable receiving device for execution by the processing device. Machine-readable media can be machine-readable storage devices, machine-readable storage substrates, random or serial access memory devices, or combinations thereof.

[0101] Various methods and formulas can be implemented in the form of computer program instructions and executed by processing devices. Suitable programming languages ​​for expressing program instructions include, but are not limited to, C, C++, variants of FORTRAN (such as FORTRAN77 or FORTRAN90), Java, Visual Basic, Perl, Tcl / Tk, JavaScript, ADA, and statistical analysis software (such as SAS, R, MATLAB, SPSS, and Stata). The various aspects of the computer-implemented methods can be written in different computational languages, and the aspects communicate with each other through appropriate system-level tools available on a given system. The processes and logic flows described in this invention can be executed by one or more programmable computers that execute one or more computer programs to perform functions by manipulating input information and generating outputs. The processes and logic flows can also be executed by special-purpose logic circuits (e.g., FPGAs (Field-Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), or RISCs), and the devices can also be implemented as special-purpose logic circuits.

[0102] A computer suitable for executing computer programs typically includes a general-purpose or special-purpose microprocessor, or both, or any other type of central processing unit. Typically, the central processing unit receives instructions and information from read-only memory or random access memory, or both. The basic components of a computer are the central processing unit for executing instructions and one or more storage devices for storing instructions and information. Typically, a computer will also include one or more mass storage devices (e.g., magneto-optical, magneto-optical, or optical discs) for storing information, or operatively coupled thereto to receive or transfer information, or both. However, a computer does not necessarily have such devices. Furthermore, a computer may be embedded in another device, such as a mobile phone, smartphone, or tablet computer, a touchscreen device or surface, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), to name just a few.

[0103] Computer-readable media suitable for storing computer program instructions and information include various forms of non-volatile memory, media, and storage devices, exemplarily including semiconductor storage devices (e.g., EPROM, EEPROM, and flash memory devices); magnetic disks (e.g., internal hard disks or removable disks); magneto-optical disks; and CD-ROMs and (Blu-ray) DVD-ROMs. Processors and memory may be supplemented or integrated therein by dedicated logic circuitry.

[0104] To provide interaction with the user, the implementation of the subject matter described in this invention can be carried out on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) for displaying information to the user) and a keyboard and pointing device (e.g., a mouse or trackball through which the user can provide input to the computer). Other types of devices can also be used to provide interaction with the user. Furthermore, the computer can interact with the user by sending and receiving documents to and from the device used by the user; for example, by sending a webpage to a web browser on the user's client device in response to a request received from a web browser.

[0105] The implementation of the subject matter described herein can be implemented in a computing system that includes backend components (e.g., as an information server), middleware components (e.g., an application server), frontend components (e.g., a client computer with a graphical user interface or web browser through which a user can interact with the implementation of the subject matter), or any combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected via digital information communication (e.g., a communication network) of any form or medium. Examples of communication networks include local area networks (“LANs”) and wide area networks (“WANs”), such as the Internet. The computing system may include clients and servers. Clients and servers are typically geographically isolated from each other and typically interact via a communication network. The client-server relationship is generated by computer programs running on their respective computers and having a client-server relationship with each other. In some embodiments, the server may be located in the cloud via a cloud computing service.

[0106] While this invention contains numerous specific implementation details, these should not be construed as limiting the scope of possible claims, but rather as descriptions of features that may be specific to a particular implementation. Certain features described in the context of a single implementation in this invention may also be implemented in combination within a single implementation. Conversely, various features described in the context of a single implementation may also be implemented separately in multiple implementations or in any suitable sub-combination. Furthermore, although features may be described above as functioning in certain combinations, or even initially claimed in this way, in some cases one or more features from the claimed combination may be removed, and the claimed combination may refer to a sub-combination or a variation of a sub-combination.

[0107] Similarly, although the operations are described in a specific order, this should not be construed as requiring such operations to be performed in the specific order shown or sequentially, or to perform all illustrated operations to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the above implementations should not be construed as requiring such separation in all implementations, and it should be understood that the described program components and systems can generally be integrated together as a single software product or packaged into multiple software products. Specific implementations of the subject matter have been described. Other implementations are within the scope of the appended claims. For example, the actions recited in the claims can be performed in different orders and still achieve the desired result. In one embodiment, the processes depicted in the figures do not necessarily require the specific order shown, or a sequential order, to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.

[0108] In one aspect, the present invention provides various computer-implemented methods described herein. In another aspect, the present invention provides one or more machine-readable hardware storage devices for the computer-implemented methods described herein, implemented by storing instructions executable by one or more data processing devices to perform the operations described herein. In another aspect, the present invention provides a system comprising: one or more data processing devices; and one or more machine-readable hardware storage devices for implementing computer-implemented methods for a test subject by storing instructions executable by the one or more data processing devices to perform the operations described herein.

[0109] Reagent test kit The present invention also provides kits for collecting, transporting, and / or analyzing samples. Such kits may contain the materials and reagents required to obtain appropriate samples (e.g., cfDNA) from subjects or to measure levels of specific biomarkers. In some embodiments, the kit includes the materials and reagents required to obtain and store samples from subjects. The samples are then transported to a service center for further processing (e.g., sequencing and / or data analysis). The kit may also include instructions for methods of collecting samples, performing tests, and interpreting and analyzing the data obtained from performing the tests.

[0110] Example The present invention is further described in the following embodiments, which do not limit the scope of the invention as described in the claims.

[0111] Example 1: Calculation of short fragment ratio based on capillary electrophoresis images In recent years, numerous studies have reported that the proportion of short cell-free DNA (cfDNA) fragments in blood / urine samples from cancer patients is significantly higher than that in blood / urine samples from healthy individuals. Currently, almost all methods used to detect the proportion of short fragments rely on next-generation sequencing (NGS), where fragment size is determined by aligning paired-end (PE) sequencing reads to the human genome and calculating their positional differences. This wet process is both complex and expensive. This paper develops a novel method for calculating the proportion of short fragments using capillary electrophoresis.

[0112] According to the article by Meng, Z. et al. ("Noninvasive detection of hepatocellular carcinoma with circulating tumor DNA features and α-fetoprotein." The Journal of Molecular Diagnostics 23.9 (2021): 1174-1184), the proportion of short fragments in a sample, such as the proportion of cfDNA fragments with a fragment size (FS) less than 150 bp (P150), is defined as the number of fragments shorter than 150 bp divided by the total number of cfDNA fragments originating from mononuclear bodies. Figure 1 As shown in A, based on NGS sequencing results, the proportion of short fragments (e.g., P150) can be determined by dividing the number of reads with fragment sizes less than 150 bp by the total number of reads with fragment sizes within the size range of a single nucleosome.

[0113] As described in this article, the start and end positions of single nuclei can also be determined by using capillary electrophoresis, and the proportion of short fragments can be calculated based on the position corresponding to 150 bp. Figure 1 B). Based on these three locations, the concentrations of fragments shorter than 150 bp (area under the curve from the start position to the 150 bp marker) and the concentrations of cfDNA fragments from mononuclear bodies (area under the curve from the start position to the end position of the mononuclear body) were calculated. The ratio between these two concentrations equals the proportion of short fragments (e.g., P150).

[0114] method Sample collection and analysis Blood collection, plasma separation, and cfDNA extraction from subjects can be performed using a standard cfDNA extraction procedure. Detailed procedures can be found in Example 1 of Chinese Patent No. 112397143B, which is incorporated herein by reference in its entirety. The extracted cfDNA was analyzed by capillary electrophoresis using an Agilent 2100 Bioanalyzer. Specifically, 1 µL of cfDNA was used with the Agilent 2100 Bioanalyzer (Agilent, model: G29939BA) and the Agilent High Sensitivity DNA Kit (Agilent, catalog number: 5067-4626) to detect cfDNA peaks and determine the concentrations (reflected by absorbance) of different fragment lengths (which migrate at different times). Determine the start position, end position, and 150 bp position. We overlay the molecular weight standard (ladder) results and sample results on the same graph. For example... Figure 2 As shown, the black line represents the electrophoresis result of a single sample, while the gray line represents the molecular weight standard result. The molecular weight standard is designed with a 150 bp molecular weight standard spot (peak), and the corresponding migration time of this peak is used as the P150 cutoff value for all samples processed in the same batch. Therefore, the area under the curve before this position represents the concentration of short fragments (e.g., fragments less than 150 bp). As another example, if it is necessary to calculate the concentration of fragments between 90 bp and 150 bp, the corresponding molecular weight standard can be prepared based on the method described above by determining the migration times corresponding to the 90 bp and 150 bp peaks. To determine the start and end positions of mononuclear bodies, the end position is defined according to the mononuclear body range (e.g., 250 bp) specified in the NGS results described above by Meng, Z. et al. A 250 bp molecular weight standard spot can be designed and prepared using the method described above. This molecular weight standard spot is used to directly determine the migration time of the 250 bp peak in the same batch of samples, thereby allowing the calculation of the mononuclear body concentration.

[0115] Furthermore, a novel method was developed to determine the initiation and termination positions of mononuclear bodies, which can be directly predicted based on the cfDNA fragment size distribution curve. According to the article by Zhu, D. et al. ("Circulating cell-free DNA fragmentation is a stepwise and conserved process linked to apoptosis." BMC Biology 21.1 (2023): 253), in the fragment size (FS) distribution of the entire cfDNA sequencing results, when the number of migrating mononuclear body cfDNA fragments tends to zero, the initiation and termination positions of mononuclear bodies can be predicted by determining the lowest point of the distribution curve. Figure 1 As shown in Figure A, these troughs appear as troughs on either side of the mononuclear body peak at approximately 167 bp. For example, since the trough between the main peaks of the first mononuclear body (167 bp) and the second mononuclear body (333 bp) is located at approximately 250 bp, the termination position of the first mononuclear body is expected to be within this range. Therefore, the termination position is determined by identifying the trough between the molecular weight standard points of 150 bp and 300 bp in the electrophoresis pattern of the sample. Figure 2 Similarly, the lowest point before 150 bp is defined as the starting position of the single-core body.

[0116] Example 2: Calculation of cfDNA concentration within the mononuclear body region based on capillary electrophoresis images Significant batch-to-batch differences were observed by comparing experimental results using different batches. Figure 3 As shown, despite using the same molecular weight standard across different batches, the migration times (X-axis) of peaks with the same cfDNA fragment size differed between batches. Therefore, it can be argued that using a fixed migration time to determine the start, end, and 150 bp positions may be problematic. Instead, it may be necessary to use the same molecular weight standard from the same batch to determine the 150 bp position and to use the lowest point within a specific range to define the start and end positions. It was also observed that the fluorescence intensity (or absorbance) (Y-axis) of a peak in the same sample varied across different batches. To address this issue, concentration calibration was performed. Specifically, the sum of the peak values ​​of the molecular weight standard peaks (e.g., 50 bp, 100 bp, 150 bp, 200 bp, and 300 bp peaks) was used as a reference for inter-batch calibration. Alternatively, a standard reference material simulating the distribution of cfDNA can be prepared, as detailed in U.S. Patent Application Publication No. 20230079748A1, which is incorporated herein by reference in its entirety. The mononuclear body concentration of this standard reference material can be used for inter-batch calibration.

[0117] Calculation of calibrated cfDNA concentration from mononuclei Due to various factors, the fragment size (FS) distribution curve of a sample often shows upward or downward shifts. In normal samples ( Figure 2 The start and end positions of the single nucleus body almost coincide with the baseline (y=0). Figure 1 The NGS sequencing plot shown in Figure A also illustrates this. However, if the curve shifts downwards, most of the mononucleosome region will lie below the baseline (y=0), leading to a significant underestimation of mononucleosome concentration when calculated directly. Figure 4 To address this issue, a fitting line can be plotted based on the start and end positions of the single nucleus body (see...). Figure 4 (The dashed line in the diagram). Next, the final cfDNA concentration from mononuclei can be calculated as the area between the mononuclei curve and this fitted line. Similarly, an upward shift of the curve may also be observed. In both cases, the start and end positions of the mononuclei can be determined first using the method described above, and then a fitted line can be drawn based on the determined start and end positions to calculate the calibrated concentration.

[0118] Optimized cfDNA concentration from mononuclear bodies for multi-cancer early detection (MCED) According to Chinese Patent No. 112397143B, there are significant differences in cfDNA concentrations in the plasma of non-cancer individuals and cancer patients. Here, we compared the concentrations of cfDNA from mononuclear bodies determined by the conventional method (v0) and the calibrated cfDNA concentrations determined by the optimized method (v2) discussed herein in a cohort (including 71 cancer samples and 120 non-cancer samples). Conventional cfDNA concentrations can be calculated using methods such as, for example, the total area under the entire FS distribution curve when using an Agilent 2100 bioanalyzer. Alternatively, as detailed in Chinese Patent No. 112397143B, cfDNA concentrations can be quantified using a Qubit™ fluorometer (Thermo Fisher Scientific, Q33226). Here, the second method described above (using a Qubit™ fluorometer) is used as a reference (cfDNA v0) for comparison with the optimized method (cfDNA v2) discussed herein.

[0119] In traditional methods, some samples may undergo hemolysis due to delayed plasma separation after blood collection or improper transport temperature. For example, when blood cells rupture, their intracellular DNA may enter the plasma, causing genomic DNA (gDNA) contamination. This also leads to an increase in the overall cfDNA concentration. Chinese Patent No. 112397143B found that the cfDNA concentration in tumor samples can be significantly higher than in normal samples, making false positives more likely. This is because gDNA fragments are always larger than 1,000 bp. Therefore, calculating the cfDNA concentration from mononuclear bodies using the method described above, instead of using the total cfDNA concentration, reduces the risk of incorporating longer DNA fragment contamination. This advantage may help eliminate false positives caused by hemolyzed samples.

[0120] like Figure 5 As shown, in a retrospective cohort study, optimized cfDNA concentration from mononuclear bodies (cfDNA v2) exhibited significantly higher values ​​in cancer patients compared to healthy controls (p = 3.4 × 10⁻⁶). -7 ).like Figure 6 As shown, the area under the curve (AUC) for cfDNA v2 is 0.721. Compared to the conventional method (using the total area under the entire FS distribution curve, cfDNA v0), the AUC increases by 0.063 when using the optimized cfDNA concentration. Notably, at high specificity levels such as 90%, the sensitivity is improved by approximately 9.6%.

[0121] Example 3: Short-fragment ratios based on capillary electrophoresis for early cancer detection The method for calculating the concentration of cfDNA fragments smaller than 150 bp was discussed in Example 2, for example, by drawing a fitted straight line between the start and end positions of the mononuclear body. Specifically, after determining the intersection of the straight line with the 150 bp marker, a line can be drawn from that intersection to the start point. The calibration method described above can be used to calculate the final area to obtain the concentration of short fragments smaller than 150 bp. Finally, the cfDNA concentration of fragments smaller than 150 bp can be divided by the cfDNA concentration from the mononuclear body to obtain the short fragment ratio (e.g., P150).

[0122] Using the same cohort as described in Example 2 to calculate the proportion of short fragments, the results showed a significant difference between cancer and non-cancer (healthy) samples, with a p-value of 6 × 10⁻⁶ for the t-test. -4 ( Figure 7 A). P150 v2, i.e., the proportion of short fragments obtained using the optimization method (cfDNA v2) discussed in this paper, has an AUC value of 0.745 ( Figure 7 B).

[0123] Example 4: Using machine learning combined with capillary electrophoresis map features for MCED Using a 90% specificity threshold to detect cancer patients, we compared the results using cfDNA concentrations from mononuclear bodies (Example 2). Figure 6 True positive cases identified by comparison with the proportion of short fragments (Example 3, Figure 7 B) Identified true positive cases. For example... Figure 8 As shown, these two methods can be complementary in cancer detection. Therefore, machine learning models, such as generalized linear models (GLM), random forests (RF), or gradient boosting machines (GBM), were employed to combine the proportion of short fragments with the concentration of cfDNA from mononuclear bodies to predict cancer probability. Specifically, a random forest (RF) model was used for this purpose.

[0124] The model can be constructed as follows. First, using the methods described in Examples 2-3, the proportion of short fragments and the concentration of cfDNA from mononuclear bodies for each sample are obtained. The proportion of short fragments and the concentration of cfDNA from mononuclear bodies for all samples are normalized using the median and MAD (median absolute difference) values ​​from healthy samples. The corrected Z-score is calculated by subtracting the healthy median from the sample value and then dividing by the MAD to account for variability. This method provides a robust measure of how much each sample deviates from the healthy baseline. Second, using the normalized proportion of short fragments and the concentration of cfDNA from mononuclear bodies as features, a random forest (RF) model is trained using 10-fold cross-validation. The average predicted from these models is defined as the final cancer probability for the test sample.

[0125] Using the same cohort as described in Example 2, the predicted probability of cancer showed a highly significant difference between cancer and non-cancer samples, with a p-value of 3.5 × 10⁻⁶ for the t-test. - ¹²( Figure 9 A). The AUC value for predicting cancer probability was 0.782 (A). Figure 9 B).

[0126] Example 5: Using machine learning combined with multidimensional cancer features for MCED In recent years, clinical practice has utilized methods for cancer screening that detect one or more protein biomarkers in the blood, such as PSA for prostate cancer screening, AFP for hepatocellular carcinoma screening, CA125 for ovarian cancer screening, CA153 for breast cancer screening, CA199 for pancreatic cancer screening, CA724 for ovarian cancer screening, CEA for gastrointestinal cancer screening, and CYFRA21-1 for breast cancer screening. Compared to other clinical tests, these methods offer advantages such as being non-invasive, automated, and cost-effective. These methods can also be used for early screening and cancer monitoring of various cancers. Further details can be found, for example, in Chinese Patent Application Publication No. 117831690A, which is incorporated herein by reference in its entirety.

[0127] By incorporating blood protein biomarkers as features into the model obtained from Example 4, superior cancer screening performance can be achieved. Furthermore, neither of these methods (regardless of whether blood protein biomarkers are included in the model) involves NGS, which can significantly reduce testing costs.

[0128] To build a model based on the model described in Example 4, seven additional protein biomarkers (AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, and CYFRA21-1), along with two features from capillary electrophoresis (cfDNA concentration and short fragment ratio from mononuclear bodies), were used as inputs to construct the cancer detection model. Following the steps described in Example 4, the model was constructed using RF methods incorporating these normalized features, and the probability of predicting cancer in the test samples was calculated. Figure 10 As shown in Figure A, the probability of predicting cancer showed a highly significant difference between cancer and non-cancer samples, with a p-value of 4.67 × 10⁻⁶ for the t-test. -22 The AUC value for predicting cancer probability was 0.890 ( Figure 10 B).

[0129] Other embodiments It should be understood that although the invention has been described in conjunction with its detailed description, the foregoing description is intended to illustrate and not limit the scope of the invention, which is defined by the scope of the appended claims. Other aspects, advantages, and modifications are within the scope of the following claims.

Claims

1. A method for determining cfDNA concentration within a specific size range, characterized in that, include: (a) Isolate cfDNA from the test sample; (b) The cfDNA is subjected to capillary electrophoresis analysis to obtain a distribution curve, wherein molecular weight standards and / or control samples are analyzed simultaneously in the same batch; the control samples contain a predetermined concentration of DNA, and the molecular weight standards contain one or more markers, and at least one marker has a predetermined DNA length; (c) Using the control sample and / or molecular weight standard as a reference, determine the concentration of cfDNA within the specified size range.

2. The method according to claim 1, characterized in that, The control sample is a circulating tumor DNA reference sample, which is prepared by inducing tumor cell apoptosis.

3. The method according to claim 1 or 2, characterized in that, The distribution curve is corrected for an overall upward or downward shift.

4. The method according to any one of claims 1-3, characterized in that, The molecular weight standard includes a first marker and a second marker, and the specific size range is predefined by the first marker and the second marker.

5. The method according to any one of claims 1-3, characterized in that, The specific size range is determined by analyzing the distribution curve.

6. The method according to claim 5, characterized in that, The specific size range is defined by the extreme points of the distribution curve, which include at least one local maximum point and / or at least one local minimum point.

7. The method according to claim 5 or 6, characterized in that, The specific size range is defined by two adjacent local minima.

8. The method according to claim 5 or 6, characterized in that, The specific size range is defined by a local maximum or a local minimum point and a pre-defined molecular weight standard marker.

9. The method according to claim 7, characterized in that, cfDNA concentration was determined by calculating the following area: The area enclosed by the distribution curve within the specified size range and the straight line connecting two adjacent local minima; Optionally, calibration is performed by measuring the corresponding area of ​​control samples and / or molecular weight standards in different batches within a reference size range, wherein the control samples and / or molecular weight standards have the same DNA concentration in each batch.

10. The method according to any one of claims 1-9, characterized in that, The specific size range corresponds to a single nucleus body.

11. A method for determining the proportion of short cfDNA fragments derived from mononuclear bodies, characterized in that, include: (a) Isolate cfDNA from the test sample; (b) The cfDNA is subjected to capillary electrophoresis analysis to obtain a distribution curve, wherein a control sample and / or a molecular weight standard are analyzed simultaneously in the same batch; the control sample contains a predetermined concentration of DNA, and the molecular weight standard contains one or more markers, and at least one marker has a predetermined DNA length. (c) Using the control sample and / or molecular weight standard as a reference, determine the proportion of short cfDNA fragments derived from mononuclear bodies.

12. The method according to claim 11, characterized in that, The control sample is a circulating tumor DNA reference sample, which is prepared by inducing tumor cell apoptosis.

13. The method according to claim 11 or 12, characterized in that, The distribution curve is corrected for an overall upward or downward shift.

14. The method according to any one of claims 11-13, characterized in that, Step (c) includes: Identify the size range of cfDNA corresponding to mononuclear bodies; Determining the first area includes: the area enclosed by the distribution curve within the specified size range and the straight line connecting two adjacent local minima; Determining the second area includes: the area jointly defined by the distribution curve, the preset molecular weight standard marker representing the upper limit of the short segment, and the straight line; Optionally, the first area and the second area are calibrated by measuring the corresponding areas of the control samples and / or molecular weight standards in different batches within a reference size range, wherein the control samples and / or molecular weight standards have the same DNA concentration in different batches; The proportion of short cfDNA fragments originating from mononuclear bodies is calculated by the ratio of the second area to the first area.

15. The method according to any one of claims 1-14, characterized in that, Further includes: Based on the concentration of cfDNA within the size range of the mononucleosome and / or the proportion of short cfDNA fragments derived from the mononucleosome, the probability that the test sample originated from a cancer patient is determined.

16. A method for predicting cancer risk, characterized in that, include: Multiple parameters were obtained from the test samples of the subjects, including: the concentration of cfDNA corresponding to the size range of mononucleosomes and the proportion of short cfDNA fragments derived from mononucleosomes; The multiple parameters are input into a machine learning system to predict cancer risk and output the prediction results. Preferably, the machine learning system is trained using a machine learning algorithm selected from Random Forest (RF), Generalized Linear Model (GLM), Support Vector Machine (SVM), and / or Gradient Boosting Machine (GBM).

17. A method for predicting cancer risk, characterized in that, include: Multiple parameters were obtained from test samples from subjects, including: the level of at least one protein biomarker, the concentration of cfDNA corresponding to the size range of a mononucleosome, and the proportion of short cfDNA fragments derived from mononucleosomes; the protein biomarkers included: AFP, CA125, CA15-3, CA19-9, CEA, CYFRA21-1, SCCA, ProGRP, and PSA; The multiple parameters are input into a machine learning system to predict cancer risk and output the prediction results. Preferably, the machine learning system is trained using a machine learning algorithm selected from Random Forest (RF), Generalized Linear Model (GLM), Support Vector Machine (SVM), and / or Gradient Boosting Machine (GBM). Optionally, the protein biomarkers include: AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA21-1, SCCA, ProGRP, and PSA.

18. The method according to claim 17, characterized in that, The plurality of parameters also includes at least one clinical parameter, which includes age, sex, and smoking status.

19. The method according to any one of claims 16-18, characterized in that, The concentration of cfDNA corresponding to the size range of mononuclear bodies and the proportion of short cfDNA fragments derived from mononuclear bodies were determined by capillary electrophoresis.

20. The method according to any one of claims 16-19, characterized in that, The concentration of cfDNA within the size range corresponding to the mononuclear body and the proportion of short cfDNA fragments derived from the mononuclear body were determined by capillary electrophoresis, including: Isolate cfDNA from the test sample; The cfDNA was analyzed by capillary electrophoresis to obtain a distribution curve, wherein molecular weight standards and / or control samples were analyzed simultaneously in the same batch; Using the control sample and / or molecular weight standard as a reference, determine the concentration of cfDNA corresponding to the size range of mononuclear bodies, and the proportion of short cfDNA fragments derived from mononuclear bodies.

21. The method according to any one of claims 16-20, characterized in that, The method can be used to assist in the early detection of at least two cancer types simultaneously.

Citation Information

Patent Citations

  • A method for predicting tumor risk scores based on plasma multi-omics multidimensional features and artificial intelligence

    CN112397143B

  • Computer implementation method for detecting abnormal signal quantification of blood sample to be detected

    CN117831690A

  • Preparation method, product, and application of circulating tumor DNA reference samples

    US20230079748A1