A blood-based multi-omics test for multi-cancer early detection
Patent Information
- Application Number
- EP2024873903
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-10
- Publication Date
- 2025-12-03
AI Technical Summary
Current cancer screening methods are inadequate for early detection due to low accuracy, invasiveness, and limited applicability to various cancer types, relying on single-dimensional biomarkers that fail to capture the multifaceted nature of cancer development.
A multi-omics approach integrating genomic, epigenetic, and proteomic features through shallow whole-genome sequencing of cell-free DNA and protein tumor markers, leveraging machine learning models to analyze copy number aberrations, fragment size, end motifs, and protein concentrations in a single blood draw for comprehensive cancer risk prediction and tissue of origin identification.
Enhances cancer detection accuracy and expands the range of detectable cancer types, enabling earlier intervention and improved patient outcomes by utilizing a multidimensional analysis of blood-based biomarkers.
Smart Images

Figure PCTCN2024087054-FTAPPB-I100001 
Figure PCTCN2024087054-FTAPPB-I100002 
Figure PCTCN2024087054-FTAPPB-I100003
Abstract
Description
A BLOOD-BASED MULTI-OMICS TEST FOR MULTI-CANCER EARLY DETECTIONTECHNICAL FIELD
[0001] The present disclosure relates to non-invasive screening and early detection of cancer using muti-omics and machine learning from a test sample. The disclosure is also related to the prediction of affected tissue of origin for those who have been predicted with high cancer risk scores.BACKGROUND
[0002] Cancer poses a serious global public health threat and is a significant health issue affecting human populations worldwide. According to the latest global cancer data released by the International Agency for Research on Cancer (IARC) of the World Health Organization in 2020, there were 19.29 million new cancer cases worldwide.
[0003] Despite advancements in the treatment of various cancers, more than two-thirds of patients are diagnosed at an advanced stage due to the lack of early symptoms. This greatly increases the difficulty of treatment, leading to high mortality rates and poor prognosis. Clinical studies have shown that the 5-year survival rate for patients with stage IA lung cancer is approximately 60%, whereas it decreases to 40%for those at stages II-IV. Early screening and detection are also crucial for the successful treatment of colorectal cancer, with 5-year survival rates exceeding 80%for stages I and II, but dropping to about 10%once distant metastasis occurs.SUMMARY
[0004] Cancer early detection could identify occult cancer lesions at the early stage of tumorigenesis when treatments are more likely to be curative and less morbid. Unfortunately, current standard-of-care (SOC) screenings are unsatisfactory in terms of accuracy, cost-effectiveness, people’s compliance, invasive as well as limited cancer types that can be screened.
[0005] Recently studies have demonstrated that blood-based cancer detection approaches may hold promise for identifying asymptomatic cancer patients from the general population. Different cancer-associated molecular characteristics, including nucleic acids, proteins, and other biomarkers circulating in blood bloodstream have been extensively exploited. However, most studies only emphasize a single dimension of analytes. For example, certain protein biomarkers, such as carcinoembryonic antigen (CEA) , prostate-specific antigen (PSA) , and cancer antigen 19-9 (CA19-9) are proven to be of clinical value in cancer monitoring, but their value in cancer early detection is still under debate due to unsatisfactory accuracy.
[0006] Tumor-derived circulating free DNA (cfDNA) has been increasingly explored since it can be extracted from the blood and tumor-specific aberrations, including copy number aberrations (CNA) , epigenetic changes, single-nucleotide mutations, cancer-derived viral sequences and chromosomal rearrangements, can be assessed with the development of sensitive techniques to provide a genomic landscape of the cancerous lesions in a patient. Although cfDNA is at the center of non-invasive cancer early detection approaches, many cfDNA-based assays only interrogated single-nucleotide mutations of cancer-related genes which may be confounded by clonal hematopoiesis of indeterminate potential (CHIP) from white blood cells (WBC) . Apart from single-nucleotide mutations, CNAs are quite common for the majority of cancer types. About 90%of solid tumors and 50%of blood-related cancers are aneuploid and have somatic CNAs. Genome-wide sequencing of cfDNA from cancer patients enabled the detection of cancer-associated copy number profiles in multiple cancer types. Moreover, the fragmentation patterns observed in an individual’s cfDNA might contain evidence of the epigenetic landscapes of the cells giving rise to these fragments. Under normal physiological conditions, the size distribution of cfDNA fragments peaks at around 167bp which bears correspondence with the chromosomal positioning of the nucleosome and its linker histone. However, nucleosome positions, patterns near transcription start sites, and the end positions of cfDNA may be altered in cancer, leading to aberrantly short DNA molecules existing in the plasma of patients with cancer.
[0007] The development of blood-based cancer detection assays by integrating a single dimension of cancer-derived biomarkers is challenging. Different cancer types acquire the functions of survival, proliferation, and dissemination via distinct mechanisms and at various stages. A great variety of molecular surrogates of cancer are shed in the bloodstream during the course of multistep tumorigenesis.
[0008] The present disclosure describes a method for cancer early detection by taking a multidimensional view of cell-free signatures that may reflect both cancer and non-cancer contributions. That incorporated both genomic, epigenetic, and proteomic features such as copy number aberration (CNA) and fragment size (FS) , end motif, oncogenic virus and protein biomarkers via shallow whole genome sequencing (sWGS) of cfDNA and measurement of the concentration of protein tumor markers (PTMs) in a single blood draw. In this application, the fragment size of the cfDNA refers to the length of a cfDNA fragment. The term “fragment size” is used interchangeably with the term “fragment length” .
[0009] The present disclosure describes methods, systems, electronic devices, and computer-readable media related to quantifying cancer signals in multiple cancer early detection (MCED) from a test blood sample, as well as methods, systems, electronic devices, and computer-readable media related to predicting tissue of origin (TOO) in subjects predicting with high cancer risk.
[0010] In one aspect, this disclosure provides a method for predicting a cancer risk for a subject. The method can be implemented by a system including one or more computers. The method can be implemented, at least in part, by a system including one or more computers. The system obtains first measurement data characterizing levels of a set of protein tumor markers in the test sample, and determines a protein-based probability of cancer (POC) metric based at least on the first measurement data. The system obtains second measurement data comprising cell-free DNA (cfDNA) sequencing data in the test sample. The determines an end motif metric characterizing end motifs of cfDNA fragments in the test sample using at least an end motif score (EMS) characterizing a degree of deviation of a distribution of the end motifs of the test sample from a control distribution. The system determines a chromosomal instability index characterizing a copy number aberration (CNA) of the cfDNA in the test sample. The system determines a fragment size (FS) metric characterizing a fragment size pattern of the cfDNA fragments in the test sample. The system processes a first model input comprising (i) the protein-based POC metric, (ii) the end motif metric, (iii) the chromosomal instability index, and (iv) the FS metric using a first machine learning model to generate prediction data characterizing at least the cancer risk for the subject. The first machine learning model has been trained through a training process using a plurality of training examples, each respective training example comprising (i) a respective training input comprising a respective set of metrics determined based on measurement data from a respective subject and (ii) a respective training output characterizing a diagnosis of the respective subject, the plurality of examples comprising (i) a set of training examples using measurement data from cancer patients and (ii) a set of training examples using measurement data from control subjects. The system outputs the prediction data characterizing the cancer risk for the subject.
[0011] In some embodiments, the panel of biomarkers are selected from at least seven different biomarkers selected from AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1, ProGRP, SCCA, and PSA.
[0012] In some embodiments, the panel of biomarkers are selected from at least 1, 2, 3 , 4, 5, 6, 7, 8, or 9 different biomarkers selected from AFP, CA125, CA15-3, CA19-9, CEA, CYFRA 21-1, ProGRP, SCCA, and PSA.
[0013] In some implementations, the panel of protein tumor markers comprises or consists of AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, and CYFRA 21-1. In some cases, the subject is a male and the panel of tumor markers comprises or consists of AFP, CA125, CA15-3, CA19-9, CEA, CYFRA 21-1, ProGRP, SCCA, and PSA. In some other cases, the subject is a female and the panel of tumor markers comprises or consists of AFP, CA125, CA15-3, CA19-9, CEA, CYFRA 21-1, ProGRP, and SCCA.
[0014] in some implementations, the system can determine the tumor metric by processing a second model input comprising the measured levels of the protein tumor markers using a second machine learning model to generate an output that specifies the protein-based POC metric.
[0015] In some cases, the second machine learning model combines two or more models selected from the group consisting of generalized linear models (GLMs) , gradient boosting machines (GBMs) , random forests (RFs) , and support vector machines (SVMs) .
[0016] For example, the second machine learning model can combine the two or more selected models using a linear combination.
[0017] In some implementations, to determine the protein-based POC metric, the system performs an outlier analysis, wherein a higher level of a protein tumor marker increases the metric.
[0018] In some implementations, to determine the FS metric, the system (i) obtains sequencing data characterizing a set of sequences of the cfDNA fragments of the test sample, (ii) determines, from the sequencing data, one or more first distribution parameters characterizing proportions of fragments within specific length ranges relative to other length ranges, (iii) determines, from the sequencing data, one or more second distribution parameters by performing binning, filtering, combining, and normalization on the sequencing data, and (iv) processes a fifth model input comprising the first distribution parameters and the second distribution parameters using a fifth machine learning model to generate an output specifying the FS metric.
[0019] In some cases, the system further determines a virus metric characterizing the abundance of one or more cancer-related viruses in the test sample. The first model input can further include the virus metric. The one or more viruses can include, for example, HBV, HPV, EBV, HHV, HIV, and / or JCPyV.
[0020] To determine the virus metric, the system obtains sequencing data characterizing sequences of the cfDNA fragments of the test sample, identifies, from the sequencing data, a first set of sequences that do not align with a human reference genome, identifies, from the first set of sequences, a second set of sequences that aligned with one of the viruses, and determines the virus metric based on a ratio of sequencing reads in the second set of sequences out of total sequencing reads.
[0021] In some implementations, to determine the end motif metric, the system obtains cfDNA sequencing data characterizing a set of cfDNA sequences of the cfDNA fragments of the test sample. For each of the set of cfDNA sequences, the system determines a respective length of the cfDNA sequence. For each of the set of cfDNA sequences, the system extracts a respective end motif as (i) a sub-sequence comprising a first predefined number of bases at a 5' end terminus of the respective cfDNA sequence or (ii) a sub-sequence comprising a second predefined number of bases before or after a breakpoint of the respective cfDNA sequence. For each of a set of different cfDNA fragment lengths, the system determines a respective distribution of different types of end motifs in the test sample.
[0022] In some implementations, the first predefined number of bases is 1bp, or 2bp, or 3bp, or 4bp, or 5bp, or 6bp, or 7bp, or 8bp. In some implementations, the second predefined number of bases is 1bp, or 2bp, or 3bp, or 4bp, or 5bp, or 6bp.
[0023] In some implementations, to determine the end motif metric, the system computes the EMS of the test sample based on (i) the distributions of the different types of end motifs for cfDNA fragments of a predefined fragment length range in the test sample and (ii) distributions of different types of end motifs for the predefined fragment length range in control samples.
[0024] In some implementations, the predefined fragment length range comprises a first length range of 50bp-250bp. In some implementations, the predefined fragment length range comprises a sub-range within a first length range of 50bp-250bp. In some implementations, the predefined fragment length range comprises a second length range of 250bp-450bp. In some implementations, the predefined fragment length range comprises a sub-range within a second length range of 250bp-450bp.
[0025] In some implementations, the method of any of claims 17-21, wherein the EMS of the test sample is computed as:
[0026] wherein i represents a type of end motif, Pi a proportion of the end motif i in all end motifs of the cfDNA fragments of the test sample, is an average or median value of the proportion of end motif i in the control samples, m is the total number of all types of end motifs.
[0027] In some implementations, the EMS of the test sample is computed as a weighted sum of (i) a first EMS determined based on a first predefined fragment length range and (ii) a second EMS determined based on a second predefined fragment length range. In some implementations, the weight coefficients of the weighted sum are determined using a third machine learning model.
[0028] In some implementations, to determine the end motif metric, the system processes a model input comprising (i) the EMS and (ii) occurrence frequencies of a selected subset of end motifs in the test sample using a fourth machine learning model to generate an output comprising the end motif metric. In some cases, the selected subset of end motifs is determined using one or more statistical tests that identify differences in end motif occurrence frequencies between healthy subjects and cancer patients.
[0029] In some implementations, before processing the first model input using the first machine learning model, the system performs the training process to train the first machine learning model.
[0030] In some implementations, in response to the cancer risk score exceeding a predefined threshold, the system predicts a tissue of origin for a cancer patient using a sixth machine learning model.
[0031] In some implementations, the system further uses the prediction data to monitor a cancer treatment of the subject.
[0032] In some implementations, to predict the cancer risk for the subject based on the test sample from the subject, the system obtains measurement data comprising cell-free DNA (cfDNA) sequencing data in the test sample, determines an end motif metric that comprises an end motif score (EMS) characterizing a degree of deviation of a distribution of end motifs of cfDNA fragments in the test sample from a control distribution, determines a cancer risk score that characterizes the cancer risk for the subject from at least the EMS, and outputs the prediction data comprising the cancer risk score for the subject.
[0033] In some cases, to determine the cancer risk score, the system processes a model input comprising (i) the EMS and (ii) occurrence frequencies of a selected subset of end motifs in the test sample using the fourth machine learning model to generate an output comprising the cancer risk score. In some cases, the selected subset of end motifs is determined using one or more statistical tests that identify differences in end motif occurrence frequencies between healthy subjects and cancer patients.
[0034] This disclosure also provides a system including one or more computers and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform the methods described above.
[0035] This disclosure also provides one or more computer storage media storing instructions that when executed by one or more computers, cause the one or more computers to perform the methods described above.
[0036] Particular embodiments of the subject matter described in this specification can be implemented so as to realize one or more of the following advantages.
[0037] Tumor markers are protein antigens or bioactive substances produced by abnormal expression of tumor cells. When tumor markers detected in blood, body fluids, or tissues reach a certain level, they can serve as indicators and signals to reflect the presence of certain tumors and the tumor burden, aiding in tumor diagnosis, differential diagnosis, prognosis assessment, and treatment monitoring. However, their sensitivity and specificity are mostly low and cannot meet clinical requirements. Moreover, the diagnosis of tumors cannot rely solely on the detection of a single tumor marker.
[0038] Copy number alterations (CNAs) are often present several years before cancer diagnosis and represent an early characteristic of cancer. For example, a study on Barrett’s esophagus, a precursor lesion to esophageal cancer, involved the analysis of copy number alterations through whole-genome low-depth sequencing of 777 biopsy specimens collected from 88 Barrett’s esophagus patients over a monitoring period spanning 15 years. The results showed that genomic copy number alterations could distinguish progressive disease (progressing to esophageal cancer) from stable disease (compared to histologically defined progression) up to 10 years prior to tissue pathology-defined transformation (Killcoyne, S., et al. Nature Medicine, 26 (11) ) .
[0039] In recent years, several studies have reported on pan-cancer early screening methods based on differential fragmentation of circulating DNA. For example, a team led by Nitzan Rosenfeld at the Cancer Research Institute of the University of Cambridge in the UK utilized the characteristic of shortened circulating DNA fragments derived from cancer. They calculated the proportion of different fragment length ranges (such as 100-150bp, 160-180bp, etc. ) and combined it with copy number alterations to establish a pan-cancer early screening model. In a study involving 200 cancer patients with 18 different cancer types and 65 healthy individuals, the model demonstrated an AUC (area under the curve) of 0.91 (Mouliere, F., et al. Science Translational Medicine, 10 (466) ) .
[0040] Liquid biopsy based on the end motif of circulating cell-free DNA (cfDNA) has shown significant potential in cancer screening. Studies have indicated that during apoptosis, DNA is cleaved by caspase-mediated endonucleases at specific sites, resulting in the release of fragmented DNA into the bloodstream, which further forms cfDNA through additional cleavage. Research has demonstrated that different endonucleases exhibit preferences for cleavage sites. For example, DNASE1L3 endonuclease tends to produce C-terminal sequence fragments, DNASE1 endonuclease tends to produce T-terminal sequence fragments, and DFFB endonuclease tends to produce A-terminal sequence fragments.
[0041] Some viruses are closely associated with the occurrence and development of cancer. For example, 80%of liver cancer patients are infected with HBV. In addition, high-risk HPV infections in women can easily lead to cervical cancer, so HPV patients need regular annual check-ups. Therefore, cancer risks can be predicted by examining the concentration of cancer-related viruses in the blood.
[0042] The present disclosure describes a method for cancer early detection by taking a multidimensional view of cell-free signatures that may reflect both cancer and non-cancer contributions. That incorporated both genomic, epigenetic, and proteomic features such as copy number aberration (CNA) and fragment size (FS) , end motif, oncogenic virus and protein biomarkers via shallow whole genome sequencing (sWGS) of cfDNA and measurement of the concentration of protein tumor markers (PTMs) in a single blood draw. The prediction system described in this specification leverages machine-learning models to perform non-invasive detection of cancer with prediction improved prediction quality.
[0043] The methods as descried herein can be used alongside with some other existing screening approaches and can offer the potential to find more types of cancer at earlier stages using one tube of blood, to improve patient outcomes by treating the disease when it is typically most responsive to therapy, and ultimately to have a profound impact on public health.
[0044] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0045] FIG. 1 shows an example schematic representation of a workflow of performing non-invasive screening and early detection of cancer from a test sample.
[0046] FIG. 2 shows a prediction system for predicting cancer risk from a test sample.
[0047] FIG. 3 is a flow diagram of an example process for predicting cancer risk from a test sample.
[0048] FIG. 4 shows the decrease of specificity in conventional clinical methods along with the number of PTMs.
[0049] FIG. 5 shows the ROC (receiver operating characteristic) curve of the cancer detection model based on the PTMs established in Example 2.
[0050] FIG. 6 shows the flowchart of the outlier analysis method incorporated into the cancer detection model based on PTMs.
[0051] FIG. 7 shows the performance comparison between the MCED model (glm POC) and optimized MCED model with outlier analysis (cutoff POC) .
[0052] FIG. 8 shows a box plot comparing the CIN score between cancer patients and the normal population in Example 5.
[0053] FIG. 9 shows the CNA analysis results of a blood plasma sample from one healthy individual, a cancer tissue sample, and a corresponding plasma sample from one cancer patient.
[0054] FIG. 10 shows the CIN score distribution in healthy and cancer pet dogs.
[0055] FIG. 11 shows a box plot comparing the FS. dist value between the cancer samples and normal samples.
[0056] FIG. 12 shows a box plot comparing the FS. dist. di of cancer samples with that of normal samples.
[0057] FIG. 13 shows a scatter plot comparing the distribution of FS. dist and FS. dist. di between cancer samples and normal samples in Example 6.
[0058] FIG. 14 shows a box plot comparing the FS. outlier score between the cancer samples and normal samples.
[0059] FIG. 15 shows the performance of different FS features (FS. dist, FS. dist. di, FS. sdsum and FS. outlier) for cancer predicting.
[0060] FIG. 16 shows the performance for cancer detection based on the machine learning model, which integrated different FS features.
[0061] FIG. 17 shows the schematic diagram of plasma cfDNA fragment end motif sequence.
[0062] FIG. 18 shows the distribution of end motif EMS score between normal individuals and cancer patients in Example 7.
[0063] FIG. 19 shows the ROC curve for the EMS score and motif diversity score (MDS) values of the end motif characteristics of cfDNA fragments.
[0064] FIG. 20 shows the ROC curve for the EMS score and MDS values of the end motif characteristics of cfDNA fragments within the first nucleosome fragment size range (50bp-250bp) in Example 7.
[0065] FIG. 21 shows the ROC curve for the EMS values and MDS values of the end motif characteristics of cfDNA fragments within the second nucleosome fragment size range (259bp-439bp) in Example 7.
[0066] FIG. 22 shows the ROC curves for the training and validation sets using EMP value in Example 8.
[0067] FIG. 23 shows the cancer detection results of each virus (HBV / HPV / EBV) reads fraction.
[0068] FIG. 24 shows the ROC curve for the linear integration model incorporating PTM (POC value) , CNA, FS, EndMotif (EMS score) , and Virus_RF calculation cancer risk score (CRS) in Example 10.
[0069] FIG. 25 shows the ROC curve for the linear integration model incorporating PTM (POC value) , CNA, FS, Endmotif (EMP value) , and Virus_RF calculation CRS in Example 10 of the present invention.
[0070] FIG. 26 shows the ROC curve for the machine learning method integrating PTM (POC value) , CNA, FS, Endmotif (EMS score) , and Virus_RF calculation CRS in Example 11.
[0071] FIG. 27 shows the ROC curve for the machine learning method integrating PTM (POC value) , CNA, FS, EndMotif (EMP value) , and Virus_RF calculation CRS in Example 11.
[0072] FIG. 28 shows the dynamic changes of CIN score before and after treatment in three lymphoma patients in Example 13.
[0073] FIG. 29 shows the dynamic changes of CRS value before and after treatment in three lymphoma patients in Example 13.
[0074] FIG. 30 is a block diagram of an example computer system.DETAILED DESCRIPTION
[0075] Example workflow
[0076] FIG. 1 shows an example workflow 100 for non-invasive screening and early detection of cancer from a test sample. As shown in FIG. 1, a test sample 115 is obtained from a subject 110. Measurement data 120 is obtained for the test sample 115. The measurement data can include, for example, protein tumor marker data and cfDNA sequencing data. Examples of data measurement techniques are described below and in the Examples.
[0077] There are many methods known in the art for measuring either gene expression (e.g., mRNA) or the resulting gene products (e.g., polypeptides or proteins) that can be used in the present methods.
[0078] In some embodiments, tumor antigen detection can be conducted using an automated immunoassay analyzer. Representative analyzers include the system from Roche Diagnostics or the Analyzer from Abbott Diagnostics. In some embodiments, the analyzers used herein include Roche cobas e411 / e601 analyzer (Roche Diagnostics GmbH, Mannheim, Germany) and Bioplex 200 platform. Using such standardized platforms permits the results from one laboratory or hospital to be transferable to other laboratories around the world. However, the methods provided herein are not limited to any one assay format or to any particular set of markers that comprise a panel.
[0079] The presence and quantification of one or more antigens or antibodies in a test sample can be determined using one or more immunoassays that are known in the art. Immunoassays typically comprise: (a) providing an antibody (or antigen) that specifically binds to the biomarker (namely, an antigen or an antibody) ; (b) contacting a test sample with the antibody or antigen; and (c) detecting the presence of a complex of the antibody bound to the antigen in the test sample or a complex of the antigen bound to the antibody in the test sample.
[0080] Well known immunological binding assays include, for example, an enzyme linked immunosorbent assay (ELISA) , which is also known as a “sandwich assay” , an enzyme immunoassay (EIA) , a radioimmunoassay (RIA) , a fluoroimmunoassay (HA) , a chemiluminescent immunoassay (CLIA) , a counting immunoassay (CIA) , a filter media enzyme immunoassay (META) , a fluorescence-linked immunosorbent assay (FLISA) , agglutination immunoassays and multiplex fluorescent immunoassays (such as the Luminex Lab MAP) , immunohistochemistry, etc.
[0081] The immunoassay can be used to determine a test amount of an antigen in a sample from a subject. First, a test amount of an antigen in a sample can be detected using the immunoassay methods described above. If an antigen is present in the sample, it will form an antibody-antigen complex with an antibody that specifically binds the antigen under suitable incubation conditions as described herein. The amount, activity, or concentration, etc. of an antibody-antigen complex can be determined by comparing the measured value to a standard or control.
[0082] Any methodology that provides for the measurement of a marker or panel of markers from a human subject is contemplated for use with the present methods. In certain embodiments, the sample from the human subject is a tissue section such as from a biopsy. In another embodiment, the sample from the human subject is a bodily fluid such as blood, serum, plasma or a part or fraction thereof. In other embodiments, the sample is a blood or serum and the markers are proteins measured therefrom. In yet another embodiment, the sample is a tissue section and the markers are mRNA expressed therein. Many other combinations of sample forms from the human subjects and the form of the markers are contemplated.
[0083] To measure tumor protein markers from a collected plasma or serum, a Roche cobas e411 / e601 analyzer (Roche Diagnostics GmbH, Mannheim, Germany) and commercial assay kits are used to detect the levels of PTMs in a total of 500 μL of plasma or serum from the blood sample, including AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1, ProGRP, SCCA, and (PSA) , following the manufacturer’s instructions.
[0084] Among these biomarkers, CA125 is a repeating peptide epitope of the mucin MUC16, which promotes cancer cell proliferation and inhibits anti-cancer immune responses. CA15-3 is derived from glycoprotein Mucin-1 (MUC-1) . CA 19-9 is a tumor-associated antigen, which was originally defined by a monoclonal antibody that has been produced by a hybridoma prepared from murine spleen cells immunized with a human colorectal cancer cell line. CA 19-9 exists in tissue as an epitope of sialyated Lewis A blood group antigen. CYFRA 21-1 is a fragment of cytokeratin 19 (KRT19) .
[0085] ProGRP is related to Gastrin-releasing peptide (GRP) . GRP is an important regulatory molecule that is implicated in a number of physiological and pathophysiological processes in humans. Its 148 amino acid preproprotein, following cleavage of a signal peptide, is further processed to produce the 27 amino acid GRP and the 68 amino acid ProGRP. Due to its short half-life of 2 minutes, it is not possible to measure GRP in blood. Therefore, an assay for the measurement of ProGRP is useful to study GRP.
[0086] SCCA include SCCA1 (also known as SERPINB3) and SCCA2 (also known as SERPINB4) . These biomarkers are known in the art, and are described e.g., Locker, et al. “ASCO 2006 update of recommendations for the use of tumor markers in gastrointestinal cancer. " Journal of clinical oncology 24.33 (2006) : 5313-5327; Del Villano BC et al: Radioimmunometric assay for a monoclonal antibody-defined tumor marker, CA 19-9. Clin Chem 29: : 549, 1983 -552; Muraro, Raffaella, et al. "Generation and characterization of B72.3 second generation monoclonal antibodies reactive with the tumor-associated glycoprotein 72 antigen. ” Cancer research 48.16 (1988) : 4588-4596; each of which is incorporated herein by reference in its entirety.
[0087] In some embodiments, a panel of biomarkers in combination with clinical parameters is selected from: AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1. In other embodiments, a panel of biomarkers is selected from AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1, ProGRP, SCCA, and PSA. In embodiments, a panel of biomarkers in combination with clinical parameters is selected from: AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1, ProGRP, SCCA, and PSA. In some embodiments, clinical parameters can be added, e.g., age, and gender. In some embodiments, the panel of biomarkers is selected from at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 different biomarkers selected from AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1, ProGRP, SCCA, and PSA. In some embodiments, the panel of biomarkers is selected from at least 7 different biomarkers selected from AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1, ProGRP, SCCA, and PSA.
[0088] In some embodiments, the panel of biomarkers comprises or consists of AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, and CYFRA 21-1. In some embodiments, the panel of biomarkers consists of AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, and CYFRA 21-1. In some embodiments, the panel of biomarkers comprises AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1, ProGRP, SCCA, and PSA. In some embodiments, the panel of biomarkers consists of AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1, ProGRP, SCCA, and PSA.
[0089] In some embodiments, the panel of biomarkers are selected from at least 1, 2, 3 , 4, 5, 6, 7, 8, or 9 different biomarkers selected from AFP, CA125, CA15-3, CA19-9, CEA, CYFRA 21-1, ProGRP, SCCA, and PSA. In some embodiments, the panel of biomarkers are selected from at least 7 different biomarkers selected from AFP, CA125, CA15-3, CA19-9, CEA, CYFRA 21-1, ProGRP, SCCA, and PSA.
[0090] In some embodiments, the subject is a male, and the panel of biomarkers comprises or consists of AFP, CA125, CA15-3, CA19-9, CEA, CYFRA 21-1, ProGRP, SCCA, and PSA. In some embodiments, the subject is a female, and the panel of biomarkers comprises or consists of AFP, CA125, CA15-3, CA19-9, CEA, CYFRA 21-1, ProGRP, and SCCA. In some embodiments, the panel of biomarkers comprises CEA, CYFRA21-1, SCCA, and ProGRP. In some embodiments, the panel of biomarkers comprise no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 different biomarkers.
[0091] In some embodiments, the panel of biomarkers consists of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 different biomarkers.
[0092] In certain embodiments, the panel of markers can comprise markers associated with a cancer selected from pancreatic cancer, ovarian cancer, liver cancer, lung cancer, stomach cancer, colorectal cancer, lymphoma, oesophageal cancer, prostate cancer, or breast cancer.
[0093] A panel can comprise any number of markers as a design choice, seeking, for example, to maximize specificity or sensitivity of the assay. Hence, an assay of interest may ask for presence of at least one of two or more biomarkers, three or more biomarkers, four or more biomarkers, five or more biomarkers, six or more biomarkers, seven or more biomarkers, eight or more biomarkers, nine or more biomarkers, ten or more biomarkers, or more as a design choice.
[0094] Isolation and extraction of cfDNA from blood (or similar fluids such as cerebrospinal fluid) can also be performed for library preparation and sequencing. The methods for extraction, library preparation, and sequencing are not specifically limited. Sequencing results can be obtained by referring to the experimental methods in the patent with the publication number CN112397143B, entitled “Method for predicting tumor risk value based on plasma multi-omics multi-dimensional features and artificial intelligence” , which is incorporated herein by reference in its entirety.
[0095] The cancer risk prediction system 200 processes the measurement data 120 to generate a prediction result 280. The prediction result 280 can include a risk score (e.g., cancer risk score) characterizing a risk for a cancer. In some implementations, the prediction result 280 can also characterize a cancer tissue of origin (TOO) . The prediction process will be described in further detail with respect to FIG. 2, and FIG. 3, and will be further described in the Examples. If the predication results suggests that the subject 110 has a high risk of cancer, further clinical diagnosis need be performed.
[0096] In some cases, a therapeutic intervention 290 can be applied to the subject 110 based on a cancer diagnosis. The therapeutic intervention 290 is an intervention of treating a cancer in the subject, reducing the rate of the increase of volume of a tumor in the subject over time, reducing the risk of developing a metastasis, and / or reducing the risk of developing an additional metastasis in the subject, upon determination of whether the subject is likely to have cancer and / or identification of the tissue of origin (TOO) in a cancer patient. In some embodiments, the treatment can halt, slow, retard, or inhibit progression of a cancer. In some embodiments, the treatment can result in the reduction of in the number, severity, and / or duration of one or more symptoms of the cancer in a subject. In some embodiments, the compositions and methods disclosed herein can be used for treatment of patients at risk for a cancer.
[0097] The treatments can generally include e.g., surgery, chemotherapy, radiation therapy, hormonal therapy, targeted therapy, and / or a combination thereof. Which treatments are used depends on the type, location and grade of the cancer as well as the patient's health and preferences. In some embodiments, the therapy is chemotherapy or chemoradiation.
[0098] In some implementations, the therapeutic intervention 290 can include administering a therapeutically effective amount of a therapeutic agent to the subject in need thereof (e.g., a subject having, or identified or diagnosed as having, a cancer) . In some embodiments, the subject has e.g., breast cancer (e.g., triple-negative breast cancer) , carcinoid cancer, cervical cancer, endometrial cancer, glioma, head and neck cancer, liver cancer, lung cancer, small cell lung cancer, lymphoma, melanoma, ovarian cancer, pancreatic cancer, prostate cancer, renal cancer, colorectal cancer, gastric cancer, testicular cancer, thyroid cancer, bladder cancer, urethral cancer, or hematologic malignancy. In some embodiments, the cancer is unresectable melanoma or metastatic melanoma, non-small cell lung carcinoma (NSCLC) , small cell lung cancer (SCLC) , bladder cancer, or metastatic hormone-refractory prostate cancer. In some embodiments, the subject has a solid tumor. In some embodiments, the cancer is squamous cell carcinoma of the head and neck (SCCHN) , renal cell carcinoma (RCC) , triple-negative breast cancer (TNBC) , or colorectal carcinoma. In some embodiments, the subject has triple-negative breast cancer (TNBC) , gastric cancer, urothelial cancer, Merkel-cell carcinoma, or head and neck cancer. In some embodiments, the subject has pancreatic cancer, ovarian cancer, liver cancer, lung cancer, stomach cancer, colorectal cancer, lymphoma, oesophageal cancer, or breast cancer.
[0099] As used herein, by an “effective amount” is meant an amount or dosage sufficient to effect beneficial or desired results including halting, slowing, retarding, or inhibiting progression of a disease, e.g., a cancer. An effective amount will vary depending upon, e.g., an age and a body weight of a subject to which the therapeutic agent is to be administered, a severity of symptoms and a route of administration, and thus administration can be determined on an individual basis. An effective amount can be administered in one or more administrations. By way of example, an effective amount is an amount sufficient to ameliorate, stop, stabilize, reverse, inhibit, slow and / or delay progression of a cancer in a patient or is an amount sufficient to ameliorate, stop, stabilize, reverse, slow and / or delay proliferation of a cell (e.g., a biopsied cell, any of the cancer cells described herein, or cell line (e.g., a cancer cell line) ) in vitro.
[0100] In some embodiments, the methods described herein can be used to monitor the progression of the disease, determine the effectiveness of the treatment, and adjust treatment strategy. For example, after the subject receives a treatment, samples can be collected from the subject. The analysis of these samples can be used to monitor the progression of the disease, determine the effectiveness of the treatment, and / or adjust treatment strategy. In some embodiments, the results are then compared to the early results. In some embodiments, a decrease of cancer risk score may suggest that the treatment is effective. In some embodiments, if the cancer risk score increases, it may suggest that the treatment is not effective. Thus, certain recommendations may be provided. These recommendations may include e.g., increasing the dosage of an existing treatment, administering one or more additional therapeutic agents to the subject, or changing therapies. In some embodiments, a significant increase of the cancer risk score indicates that the subject is at high risk (e.g., the cancer or some other underlying disease progresses very rapidly) , and this subject should be closely monitored (e.g., being admitted to hospital or intensive care unit) .
[0101] In some embodiments, the methods described herein can be used to monitor the progression of the disease. In some embodiments, samples are collected from the subject prior to receiving the treatment. Cancer risk score is determined based on the methods as described herein. The subject then receives the treatment. After the treatment, samples are collected from the subject after the initial treatment. Cancer risk score is then determined based on the methods as described herein. If the subject receives the treatment for a period of time (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12months) , samples can be collected from the subject during the treatment period and / or after the treatment. Cancer risk score can be determined. The change of cancer risk score can be tracked.
[0102] The treatments can generally include e.g., surgery, chemotherapy, radiation therapy, hormonal therapy, targeted therapy, and / or a combination thereof. In some embodiments, the therapeutic agent can comprise one or more inhibitors selected from the group consisting of an inhibitor of B-Raf, an EGFR inhibitor, an inhibitor of a MEK, an inhibitor of ERK, an inhibitor of K-Ras, an inhibitor of c-Met, an inhibitor of anaplastic lymphoma kinase (ALK) , an inhibitor of a phosphatidylinositol 3-kinase (PI3K) , an inhibitor of an Akt, an inhibitor of mTOR, a dual PI3K / mTOR inhibitor, an inhibitor of Bruton's tyrosine kinase (BTK) , and an inhibitor of Isocitrate dehydrogenase 1 (IDH1) and / or Isocitrate dehydrogenase 2 (IDH2) . In some embodiments, the additional therapeutic agent is an inhibitor of indoleamine 2, 3-dioxygenase-1) (IDO1) (e.g., epacadostat) .
[0103] In some embodiments, the therapeutic agent can comprise one or more inhibitors selected from the group consisting of an inhibitor of HER3, an inhibitor of LSD1, an inhibitor of MDM2, an inhibitor of BCL2, an inhibitor of CHK1, an inhibitor of activated hedgehog signaling pathway, and an agent that selectively degrades the estrogen receptor.
[0104] In some embodiments, the therapeutic agent can comprise one or more therapeutic agents selected from the group consisting of Trabectedin, nab-paclitaxel, Trebananib, Pazopanib, Cediranib, Palbociclib, everolimus, fluoropyrimidine, IFL, regorafenib, Reolysin, Alimta, Zykadia, Sutent, temsirolimus, axitinib, everolimus, sorafenib, Votrient, Pazopanib, IMA-901, AGS-003, cabozantinib, Vinflunine, an Hsp90 inhibitor, Ad-GM-CSF, Temazolomide, IL-2, IFNa, vinblastine, Thalomid, dacarbazine, cyclophosphamide, lenalidomide, azacytidine, lenalidomide, bortezomid, amrubicine, carfilzomib, pralatrexate, and enzastaurin.
[0105] In some embodiments, the therapeutic agent can comprise one or more therapeutic agents selected from the group consisting of an adjuvant, a TLR agonist, tumor necrosis factor (TNF) alpha, IL-1, HMGB1, an IL-10 antagonist, an IL-4 antagonist, an IL-13 antagonist, an IL-17 antagonist, an HVEM antagonist, an ICOS agonist, a treatment targeting CX3CL1, a treatment targeting CXCL9, a treatment targeting CXCL10, a treatment targeting CCL5, an LFA-1 agonist, an ICAM1 agonist, and a Selectin agonist.
[0106] In some embodiments, carboplatin, nab-paclitaxel, paclitaxel, cisplatin, pemetrexed, gemcitabine, FOLFOX, or FOLFIRI are administered to the subject.
[0107] In some embodiments, the therapeutic agent is an antibody or antigen-binding fragment thereof. In some embodiments, the therapeutic agent is an antibody that specifically binds to PD-1, CTLA-4, BTLA, PD-L1, CD27, CD28, CD40, CD47, CD137, CD154, TIGIT, TIM-3, GITR, or OX40.
[0108] In some embodiments, the therapeutic agent is an anti-PD-1 antibody, an anti-OX40 antibody, an anti-PD-L1 antibody, an anti-PD-L2 antibody, an anti-LAG-3 antibody, an anti-TIGIT antibody, an anti-BTLA antibody, an anti-CTLA-4 antibody, or an anti-GITR antibody.
[0109] In some embodiments, the therapeutic agent is an anti-CTLA4 antibody (e.g., ipilimumab) , an anti-CD20 antibody (e.g., rituximab) , an anti-EGFR antibody (e.g., cetuximab) , an anti-CD319 antibody (e.g., elotuzumab) , or an anti-PD1 antibody (e.g., nivolumab) .
[0110] Example prediction system
[0111] FIG. 2 shows an example prediction system 200 for predicting a cancer risk for a subject based on a test sample from the subject. The prediction system 200 is an example of a system implemented as computer programs on one or more computers in one or more locations in which the systems, components, and techniques described below are implemented.
[0112] The prediction system 200 is configured to obtain input data 210 including one or more of: protein tumor marker data, cfDNA sequencing data, and / or clinical data of the subject. The prediction system 200 uses one or more metric prediction models 220 to predict metrics data for generating a model input 230. The model input 230 can include a protein-based POC metric, an end motif metric, a chromosomal instability index, an FS metric, and / or a virus metric.
[0113] The prediction system 200 uses a prediction machine learning model 240 (e.g., a first machine learning model) to process the model input 230 to generate the prediction results 280. The first machine learning model has been trained through a training process using a plurality of training examples, each respective training example comprising (i) a respective training input comprising a respective set of metrics determined based on measurement data from a respective subject and (ii) a respective training output characterizing a diagnosis of the respective subject, the plurality of examples comprising (i) a set of training examples using measurement data from cancer patients and (ii) a set of training examples using measurement data from control subjects. The prediction machine-learning model 240 can be any appropriate type of machine-learning model.
[0114] Example prediction process
[0115] FIG. 3 is a flow diagram of an example process 300 for predicting cancer risk from a test sample. For convenience, the process 300 will be described as being performed by a system of one or more computers located in one or more locations. For example, a prediction system, e.g., the prediction system 200 of FIG. 2, appropriately programmed in accordance with this specification, can perform the process 300.
[0116] At 310, the system obtains first measurement data characterizing levels of a set of protein tumor markers in the test sample. In some cases, the set of protein tumor markers includes at least seven, eight, nine, or ten protein tumor markers selected from AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1, ProGRP, SCCA, and PSA. In some cases, the set of protein tumor markers includes AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, and CYFRA 21-1. In some implementations, the set of protein tumor markers includes AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1, ProGRP, SCCA, and PSA. In some implementations, the set of protein tumor markers includes AFP, CA125, CA15-3, CA19-9, CEA, CYFRA 21-1, ProGRP, SCCA, and PSA.
[0117] At 320, the system determines a protein-based POC metric based at least on the first measurement data. In some implementations, to determine the protein-based POC metric, the system performs an outlier analysis, wherein a higher level of a protein tumor marker increases the tumor metric.
[0118] The system can determine the tumor metric by processing a second model input comprising the measured levels of the protein tumor markers using a second machine learning model to generate an output that specifies the tumor metric.
[0119] In some cases, the second machine learning model combines two or more models selected from the group consisting of generalized linear models (GLMs) , gradient boosting machines (GBMs) , random forests (RFs) , and support vector machines (SVMs) .
[0120] For example, the second machine learning model can combine the two or more selected models using a linear combination.
[0121] At 330, the system obtains second measurement data comprising cell-free DNA (cfDNA) sequencing data in the test sample.
[0122] At 340, the system determines an end motif metric characterizing end motifs of cfDNA fragments in the test sample using at least an end motif score (EMS) characterizing a degree of deviation of a distribution of the end motifs of the test sample from a control distribution.
[0123] For example, to determine the end motif metric, the system obtains cfDNA sequencing data characterizing a set of cfDNA sequences of the cfDNA fragments of the test sample. For each of the set of cfDNA sequences, the system determines a respective length of the cfDNA sequence. For each of the set of cfDNA sequences, the system extracts a respective end motif as (i) a sub-sequence comprising a first predefined number of bases at a 5' end terminus of the respective cfDNA sequence or (ii) a sub-sequence comprising a second predefined number of bases before or after a breakpoint of the respective cfDNA sequence. For each of a set of different cfDNA fragment lengths, the system determines a respective distribution of different types of end motifs in the test sample.
[0124] In some implementations, the first predefined number of bases is 1bp, or 2bp, or 3bp, or 4bp, or 5bp, or 6bp, or 7bp, or 8bp. In some implementations, the second predefined number of bases is 1bp, or 2bp, or 3bp, or 4bp, or 5bp, or 6bp.
[0125] In some implementations, to determine the end motif metric, the system computes the EMS of the test sample based on (i) the distributions of the different types of end motifs for cfDNA fragments of a predefined fragment length range in the test sample and (ii) distributions of different types of end motifs for the predefined fragment length range in control samples.
[0126] In some implementations, the predefined fragment length range comprises a first length range of 50bp-250bp. In some implementations, the predefined fragment length range comprises a sub-range within a first length range of 50bp-250bp. In some implementations, the predefined fragment length range comprises a second length range of 250bp-450bp. In some implementations, the predefined fragment length range comprises a sub-range within a second length range of 250bp-450bp.
[0127] In some implementations, the EMS of the test sample is computed as:
[0128] wherein i represents a type of end motif, Pi a proportion of the end motif i in all end motifs of the cfDNA fragments of the test sample, is an average or median value of the proportion of end motif i in the control samples, m is the total number of all types of end motifs.
[0129] In some implementations, the EMS of the test sample is computed as a weighted sum of (i) a first EMS determined based on a first predefined fragment length range and (ii) a second EMS determined based on a second predefined fragment length range. In some implementations, the weight coefficients of the weighted sum are determined using a third machine learning model.
[0130] At 350, the system determines a chromosomal instability index characterizing a copy number aberration (CNA) of the cfDNA fragments in the test sample; and
[0131] At 360, the system determines a fragment size (FS) metric characterizing a fragment size pattern of the cfDNA fragments in the test sample.
[0132] In some implementations, to determine the FS metric, the system (i) obtains sequencing data characterizing a set of sequences of the cfDNA fragments of the test sample, (ii) determines, from the sequencing data, one or more first distribution parameters characterizing proportions of fragments within specific length ranges relative to other length ranges, (iii) determines, from the sequencing data, one or more second distribution parameters by performing binning, filtering, combining, and normalization on the sequencing data, and (iv) processes a fifth model input comprising the first distribution parameters and the second distribution parameters using a fifth machine learning model to generate an output specifying the fragment metric.
[0133] At 370, the system processes a first model input comprising (i) the protein-based POC metric, (ii) the end motif metric, (iii) the chromosomal instability index, and (iv) the FS metric using a first machine learning model to generate prediction data characterizing at least the cancer risk for the subject. The first machine learning model has been trained through a training process using a plurality of training examples, each respective training example comprising (i) a respective training input comprising a respective set of metrics determined based on measurement data from a respective subject and (ii) a respective training output characterizing a diagnosis of the respective subject, the plurality of examples comprising (i) a set of training examples using measurement data from cancer patients and (ii) a set of training examples using measurement data from control subjects.
[0134] At 380, the system outputs the prediction data characterizing the cancer risk for the subject.
[0135] In some cases, the system further determines a virus metric characterizing abundances of one or more cancer-related viruses in the test sample. The first model input can further include the virus metric. The one or more viruses can include, for example, HBV, HPV, EBV, HHV, HIV, and / or JCPyV.
[0136] To determine the virus metric, the system obtains sequencing data characterizing sequences of the cfDNA fragments of the test sample, identifies, from the sequencing data, a first set of sequences that do not align with a human reference genome, identifies, from the first set of sequences, a second set of sequences that aligned with one of the viruses, and determines the virus metric based on a ratio of sequencing reads in the second set of sequences out of total sequencing reads.
[0137] In some implementations, before processing the first model input using the first machine learning model, the system performs the training process to train the first machine learning model.
[0138] In some implementations, in response to the cancer risk score exceeding a predefined threshold, the system further predicts a tissue of origin for a cancer patient using a sixth machine learning model.
[0139] In some implementations, the system further uses the prediction data to monitor a cancer treatment of the subject.
[0140] Kits
[0141] The present disclosure also provides kits for collecting, transporting, and / or analyzing samples. Such a kit can include materials and reagents required for obtaining an appropriate sample from a subject, or for measuring the levels of particular biomarkers. In some embodiments, the kits include those materials and reagents that would be required for obtaining and storing a sample from a subject. The sample is then shipped to a service center for further processing (e.g., sequencing and / or data analysis) .
[0142] The kits may further include instructions for collect the samples, performing the assay and methods for interpreting and analyzing the data resulting from the performance of the assay.
[0143] The invention is further described in the following examples, which do not limit the scope of the invention described in the claims.
[0144] Computer system
[0145] FIG. 30 is a block diagram of an example computer system 800 that can be used to perform operations described above. The system 800 includes a processor 810, a memory 820, a storage device 830, and an input / output device 840. Each of the components 810, 820, 830, and 840 can be interconnected, for example, using a system bus 850. The processor 810 is capable of processing instructions for execution within the system 800. In one implementation, the processor 810 is a single- threaded processor. In another implementation, the processor 810 is a multi-threaded processor. The processor 810 is capable of processing instructions stored in the memory 820 or on the storage device 830.
[0146] The memory 820 stores information within the system 800. In one implementation, the memory 820 is a computer-readable medium. In one implementation, the memory 820 is a volatile memory unit. In another implementation, the memory 820 is a non-volatile memory unit.
[0147] The storage device 830 is capable of providing mass storage for the system 800. In one implementation, the storage device 830 is a computer-readable medium. In various different implementations, the storage device 830 can include, for example, a hard disk device, an optical disk device, a storage device that is shared over a network by multiple computing devices (for example, a cloud storage device) , or some other large capacity storage device.
[0148] The input / output device 840 provides input / output operations for the system 800. In one implementation, the input / output device 840 can include one or more network interface devices, for example, an Ethernet card, a serial communication device, for example, a RS-232 port, and / or a wireless interface device, for example, a 502.11 card. In another implementation, the input / output device can include driver devices configured to receive data and send output data to other input / output devices, for example, keyboard, printer and display devices 860. Other implementations, however, can also be used, such as mobile computing devices, mobile communication devices, set-top box television client devices, etc.
[0149] Although an example processing system has been described in FIG. 30, implementations of the subject matter and the functional operations described in this disclosure can be implemented in other types of digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this disclosure and their structural equivalents, or in combinations of one or more of them.
[0150] This specification uses the term “configured” in connection with systems and computer program components. For a system of one or more computers to be configured to perform particular operations or actions means that the system has installed on it software, firmware, hardware, or a combination of them that in operation cause the system to perform the operations or actions. For one or more computer programs to be configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by data processing apparatus, cause the apparatus to perform the operations or actions.
[0151] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or in addition, the program instructions can be encoded on an artificially-generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus.
[0152] The term “data processing apparatus” refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit) . The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0153] A computer program, which may also be referred to or described as a program, software, a software application, an app, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub-programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communication network.
[0154] In this specification the term “engine” is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components, installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and running on the same computer or computers.
[0155] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA or an ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers.
[0156] Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA) , a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few.
[0157] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0158] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s device in response to requests received from the web browser. Also, a computer can interact with a user by sending text messages or other forms of message to a personal device, e.g., a smartphone that is running a messaging application, and receiving responsive messages from the user in return.
[0159] Data processing apparatus for implementing machine-learning models can also include, for example, special-purpose hardware accelerator units for processing common and compute-intensive parts of machine learning training or production, i.e., inference, workloads.
[0160] Machine learning models can be implemented and deployed using a machine learning framework, e.g., a TensorFlow framework.
[0161] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back-end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front-end component, e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN) , e.g., the Internet.
[0162] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some implementations, a server transmits data, e.g., an HTML page, to a user device, e.g., for purposes of displaying data to and receiving user input from a user interacting with the device, which acts as a client. Data generated at the user device, e.g., a result of the user interaction, can be received at the server from the device.
[0163] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
[0164] Similarly, while operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0165] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.
[0166] EXAMPLES
[0167] The invention is further described in the following examples, which do not limit the scope of the invention described in the claims.
[0168] Example 1
[0169] In the first example, a non-invasive liquid biopsy technique is described. The non-invasive liquid biopsy technique allows the isolation and extraction of cfDNA from blood (or similar fluids such as cerebrospinal fluid) for library preparation and sequencing. The methods for extraction, library preparation, and sequencing are not specifically limited. Sequencing results can be obtained by referring to the experimental methods in the patent with the publication number CN112397143B, entitled “Method for predicting tumor risk value based on plasma multi-omics multi-dimensional features and artificial intelligence, ” which is incorporated herein by reference in its entirety.
[0170] Example 2: Protein quantification and analysis
[0171] For the collected plasma or serum, a Roche cobas e411 / e601 analyzer (Roche Diagnostics GmbH, Mannheim, Germany) and commercial assay kits are used to detect the levels of PTMs in a total of 500 μL of plasma or serum from each blood sample, including AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1, ProGRP, SCCA, and PSA, following the manufacturer’s instructions.
[0172] 591 cancer patients and 1055 non-cancer individuals were recruited as the training set for the detection (developed by the SeekIn laboratory) . Additionally, 363 cancer patients and 5556 non-cancer individuals from a collaborating laboratory were collected to form an independent validation set 1. The detection data has been anonymized. Eligible cancer patients had been pathologically diagnosed and had not received treatment before the blood draw. Non-cancer individuals had no history of cancer. The study analyzes clinical data and PTM quantification data from 1005 cancer patients and 812 non-cancer individuals previously published by Cohen et al., “Detection and localization of surgically resectable cancers with a multi-analyte blood test, ” Science, 359 (6378) , 926–930. Q1. (2018) as an independent validation set 2. The published data from Johns Hopkins University only includes six tumor markers and lacks the value for CA72-4. The inventors used the average value of CA72-4 from the non-cancer sample group in the training set to replace the missing values.
[0173] Almost all PTMs in the training set exhibit a cancer detection specificity exceeding 95.0%, with the exception of CA72-4 at 85.7%. However, the individual sensitivity of single PTMs for detecting malignant tumors is very low (ranging from 0%to 70.6%, with a median of 16.1%) . Furthermore, when analyzing a single tumor marker, the specificity is high but the sensitivity is low. According to traditional clinical methods based on PTM quantification, false positive rates accumulate when analyzing multiple tumor markers, as the sensitivity is low and specificity decreases with an increasing number of PTMs (see FIG. 4) . The traditional clinical method mentioned here is based on PTM quantification and assesses results using a single threshold based on the manufacturer’s recommended predetermined reference ranges for each PTM (i.e., manufacturer-recommended cutoff values: AFP at 5.8 IU / ml, CA125 at 35.0 U / ml, CA15-3 at 26.4 U / ml, CA19-9 at 27.0 U / ml, CA72-4 at 6.9 U / ml, CEA at 4.7 ng / ml, CYFRA21-1 at 3.3 ng / ml) . As per the data shown in Table 1, simultaneous detection of these seven PTMs in all samples yields a specificity of only 69.4% (95%CI: 66.5%to 72.2%) . In summary, these results indicate that the traditional clinical approach utilizing combinations of multiple protein tumor markers exhibits a significantly high false positive rate.
[0174] Table 1: Performance of MCED comparison between clinical method and machine learning method
[0175] *Individuals with at least one of the markers included in the panel showing values above the cut-off point were considered to be positive. PPV, positive predictive value. NPV, negative predictive value.
[0176] Establishing an early cancer screening model using artificial intelligence (AI) methods
[0177] Considering the high false positive rate of traditional clinical methods, a highly specific and robust Multi-Cancer Early Detection (MCED) method is crucial in the clinical setting. An algorithm has been developed using artificial intelligence, based on the expression of seven PTMs and individuals’ clinical information including gender and age, to differentiate between cancer and non-cancer individuals by calculating a cancer signal index (POC) . Common machine learning algorithms such as GLM, GBM, RF, and SVM were considered (this Example used the GLM method) to establish the MCED model. The model was able to distinguish between cancer and non-cancer controls in the training set (AUC = 0.868) as well as in two independent validation sets (AUC = 0.744 and 0.818 respectively; see FIG. 5) . With a specificity of approximately 90.0%, the sensitivity in these three groups was 58.2% (training set, 95% CI: 54.1%to 62.2%) , 47.4% (independent validation set 1, 95%CI: 42.1%to 52.7%) , and 49.3% (independent validation set 2, 95%CI: 46.1%to 52.4%) .
[0178] In particular, an optimized Multi-Cancer Early Detection (MCED) model has been developed for the early detection of multiple cancers.
[0179] The MCED model has been trained using a training set comprising 1646 samples, each including the expression levels of 7 selected protein tumor markers (AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1) and clinical information such as age and gender. Initially, the inventors standardized the data using the Z-score method, a commonly used approach for normalizing data to have a mean of 0 and a standard deviation of 1. However, if the expression levels of a protein tumor marker do not follow a Gaussian distribution, for example, exhibiting skewness or containing outliers, modifying the Z-score can better handle these non-normally distributed data by computing the difference between the observed value and the median, and then dividing by the median absolute deviation (MAD) .
[0180] After standardization, these datasets were used to develop the MCED model using machine learning methods such as Gradient Boosting Machine (GBM) , Generalized Linear Model (GLM) , Random Forest (RF) , and Support Vector Machine (SVM) . These machine learning methods were employed to construct different models capable of distinguishing between cancer and non-cancer samples.
[0181] In order to optimize model performance, 10-fold cross-validation was performed for each machine learning method. The cross-validation process generated average cancer prediction scores ranging from 0 to 1 for each machine learning method. These average cancer prediction scores from all four machine learning methods were combined into a matrix and used as input for the second-layer GLM algorithm. The GLM algorithm integrated these scores from different machine learning methods and constructed the final ensemble model, generating the final cancer signal scores, referred to as the Possibility of Cancer (POC) index, for each test sample. A higher POC index indicates a higher cancer signal score for the test sample. The performance and robustness of the MCED model were validated using two independent sets.
[0182] This comprehensive approach aims to combine models trained using different methods. Model integration can capture different aspects of the data, reduce the risk of overfitting, and enhance the accuracy and reliability of the MCED model in determining whether a sample is cancerous. Compared to the aforementioned method, the optimized MCED model can detect an additional 18 cancer samples. With the same specificity of 90.0%, the sensitivity has increased to 61.3%, representing a 3.1%improvement in sensitivity.
[0183] Example 3: Combining the optimized MCED model with outlier analysis
[0184] Certain protein tumor markers that are associated with multiple cancer types contribute more or have higher weights in the MCED model. Conversely, protein tumor markers that are highly specific for particular cancer types (e.g., AFP specifically used for liver cancer detection) may have relatively lower contributions. This leads to a scenario where a highly specific protein tumor marker for a particular cancer type shows an abnormally high level, while other protein tumor markers remain normal, the MCED model may predict a lower POC.
[0185] To address this issue, an outlier analysis method has been developed as a supplement to predict these types of cancer samples (as shown in FIG. 6) . The outlier analysis method focuses on identifying and analyzing cases where highly specific cancer protein tumor markers exhibit abnormal expression levels compared to normal cases. Incorporating this method into the MCED model, improves the detection of tumor samples that may present abnormal protein tumor marker expression. Here, the inventors based on over 6000 non-cancer samples used the following three methods to determine the cutoff value for outlier analysis.
[0186] Boxplot Method: The boxplot method identifies outliers by plotting boxplots of protein tumor markers in the normal control samples. The boxplot displays the interquartile range, and observations exceeding 1.5 times the interquartile range above the upper quartile can be considered as the cutoff value for outliers.
[0187] Modified Z-score: Some non-cancerous diseases can also result in elevated expression levels of protein tumor markers in the normal control group, leading to skewness in the protein expression levels in the normal control group. Therefore, using the difference between the observation and the median divided by the median absolute deviation, the modified Z-score is considered to address data skewness. The expression of protein tumor markers with a modified Z-score > 10 is defined as the cutoff value for outliers.
[0188] Percentile Method: The percentile method compares the observations with the percentiles of the data, and observations exceeding the 99th percentile of the normal control group are considered the cutoff value for outliers.
[0189] Based on the cutoff values obtained from the three methods mentioned above, the maximum value is selected as the final cutoff value for high outliers. If the expression level of a specific protein tumor marker in the test sample exceeds the corresponding outlier cutoff value, then the sample can be predicted to carry a cancer signal. By developing this outlier analysis method, the inventors aim to enhance the identification of cancer samples with significantly abnormal expression levels of a cancer-specific protein tumor marker. By effectively predicting these specific samples, the inventors can improve the sensitivity for detecting specific types of cancer.
[0190] For example, alpha-fetoprotein (AFP) is a specific protein tumor marker for liver cancer and has a relatively lower weight in the inventors’ multi-cancer early detection model. However, by integrating the outlier analysis described above, the inventors were able to improve the sensitivity of liver cancer detection. Through outlier analysis, the outlier cutoff value for AFP was determined to be 614.4 IU / ml. Solely based on the MCED model, 154 out of 244 liver cancer samples were successfully predicted as positive, with a specificity of 92.9% (95%confidence interval: 92.3%to 93.5%) and a sensitivity of 63.1% (95%confidence interval: 56.7%to 69.2%) . By combining outlier analysis with the optimized MCED model, an additional 23 positive samples were successfully predicted, resulting in a total of 177 positive samples. The sensitivity increased to 72.5% (95%confidence interval: 66.5%to 78.0%) , representing a 9.4%improvement, while the specificity also slightly increased to 93.4%. As shown in FIG. 7, the AUC increased from 0.865 to 0.907, and the Delong test demonstrated a significant increase in the AUC value (P < 0.001) .
[0191] Example 4: The quantified results of 10 protein markers are used for cancer screening
[0192] PTM quantification: Conventional venipuncture is performed to collect blood, and within the specified timeframes, plasma or serum is obtained through separation. The quantification of protein tumor markers is conducted using standard methods. In this Example, electrochemiluminescence immunoassay (Roche cobas e411 / e601) and commercially available assay kits compatible with this platform are utilized (e.g., for the detection of SCCA levels, the Roche Elecsys SCC immunoassay kit is employed) . Following the manufacturer’s instructions, the expression levels of 10 PTMs in the test samples are determined, including AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1, ProGRP, SCCA, and PSA.
[0193] Quantitative data was obtained for protein markers from 706 samples using the mentioned method, comprising 203 male and 503 female samples. Each male sample includes quantification of 10 selected protein tumor markers (AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1, ProGRP, SCCA, PSA) , while each female sample includes the expression levels of 9 selected protein tumor markers (AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1, ProGRP, SCCA) , as well as clinical information such as age and gender. Since PSA is a prostate-specific antigen, it is not required for female samples. Use the method from Example 2 establishing a model using the GLM algorithm, and combining it with the outlier analysis method from Example 5 to determine the cutoff values for abnormal protein tumor markers, the sensitivity of the model for multi-cancer early screening has been further enhanced.
[0194] The model combining the methods for the 10 markers with the cutoff values for outliers achieved an AUC of 0.797 for cancer detection in the aforementioned samples, with a specificity of 89.9% (95%confidence interval: 84.1%to 94.1%) and a sensitivity of 55.1% (95%confidence interval: 50.8%to 59.3%) . For comparison, using only 7 protein tumor markers (AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1) on the same dataset yielded an AUC of 0.786, with a specificity of 89.9%(95%confidence interval: 84.1%to 94.1%) and a sensitivity of 52.9% (95%confidence interval: 48.6%to 57.2%) . The ROC curves of the two models showed significant differences (DeLong’s test: p-value =0.02) . As ProGRP, SCCA, and PSA are highly specific to small cell lung cancer (SCLC) , squamous cell carcinoma, and prostate cancer, respectively, the performance of the 10 protein tumor markers model significantly improved for these cancer types. The detection rate in small cell lung cancer samples was 90.9% (20 / 22) , representing a 40.9%improvement compared to the 7 protein tumor markers model; for cervical cancer samples, the detection rate was 50% (4 / 8) , indicating a 37.5%improvement; and for prostate cancer samples, the detection rate was 85.7% (6 / 7) , showing a 42.9%improvement.
[0195] Example 5. CNA analysis
[0196] Based on the method of Example 1, 1197 samples were collected (617 cancer samples and 580 healthy samples) , and library sequencing was performed to obtain the raw data. Subsequently, the raw sequencing data was aligned to the human reference genome. After that, the human reference genome was divided into non-overlapping bins of the same size (e.g., 100kb, 200kb, 500kb) , and the number of reads in each bin was counted. Additionally, based on factors such as the GC content and the proportion of N in each bin, bins with high GC content and abnormally high N content were filtered out. Furthermore, variations in the same bins in a control group (e.g., normal blood samples) were used to filter the extent of variation, for example, by filtering bins with abnormal standard deviation in the normal group or by filtering out the highest 5%or 10%of bins based on standard deviation. Following this filtering process, GC correction and N proportion correction were performed on the filtered bins.
[0197] To reduce the impact of background noise and improve the signal-to-noise ratio, adjacent bins can be combined into larger bins (e.g., 1M, 5M) . In some cases, these combined bins can even represent the long and short arms of chromosomes, such as 1q, 1p, 2q, 2p, 3q, 3p, and so on. Based on normal control samples, the distribution of reads on the same synthesized larger bins follows a normal distribution. This allows for the calculation of the average number of reads and standard deviation (SD) for each bin, enabling Z-score standardization for each bin. By setting a cutoff value for the Z-score, such as 3, any Z-score greater than 3 or less than -3 indicates an abnormal region, and the number of abnormal bins in each sample is defined as the CIN score. It has been observed that the number of abnormal bins in cancer patients is significantly higher than in normal individuals. As shown in FIG. 8, the number of abnormal bins in cancer patients is significantly greater than in normal individuals.
[0198] In the previous step, after calculating the Z-score for each bin, the “Circular Binary Segmentation” algorithm or hidden Markov model can be used to combine 1Mb bins into different CNA segments. The average Z-score for each segment is calculated. If there are abnormal Z-score segments present, it could indicate that the sample may be from a cancer patient. The more abnormal segments there are, the more significant the abnormality, and the higher the likelihood that it originates from a tumor. FIG. 9 shows the CNA analysis results for a healthy individual (A) , tumor tissue from a cancer patient (B) , and plasma (C) .
[0199] Using the above method, the inventors analyzed plasma samples from 3 pet dogs with malignant tumors and 1 pet dog with a benign tumor. The inventors compared these samples with blood samples from 38 healthy pet dogs (referencing our study “Circulating cell-free DNA fragmentation is a stepwise and conserved process linked to apoptosis. BMC Biol. 2023; 21 (1) : 253. ” ) . The CIN scores were compared (FIG. 10) , revealing that the CIN score for the 3 pet dogs with malignant tumors was the highest. In contrast, the pet dog with the benign tumor had a CIN score within the range of the healthy dogs. Therefore, CNA analysis can be used for cancer detection in pet dogs.
[0200] Example 6: Using FS analysis to compute the probability of the sample origination
[0201] Based on the sequencing alignment results, filtering out low-quality aligned reads and retaining high-quality alignments (>30) , the distribution of insert fragment size (the distance between the two ends of reads aligned to the chromosome) is calculated. Based on the probability of cancer origin from the primary nucleosome, the inventors computed the difference between the proportion of insert fragments between 80-150bp and those between 30-250bp, then subtracted the proportion between 175-220bp, denoted as FS. dist. Using the same dataset as in Example 5, as shown in FIG. 11, the difference in FS. dist between cancer and normal samples indicates a significant distinction between cancer and normal samples.
[0202] As shown in Table 2, the majority of samples (>90%) exhibited a main peak position in the fragment size histogram distribution at either 166 or 167 bp. Based on the peak insert fragment size of each sample, they were divided into two groups: those with peaks less than or equal to 166 and those with peaks greater than or equal to 167. The FS. dist for each sample was then calculated. It was observed that within the same group, the FS. dist of cancer samples generally exceeded that of normal samples.
[0203] Table 2: FS examples
[0204] FIG. 11 displays a box plot of FS. dist in cancer patients and healthy individuals, demonstrating a significant difference in FS. dist between cancer patients and normal individuals. Based on the probability of cancer origin from the secondary nucleosome, The inventors calculated the proportion of fragments between 260 and 325 bp out of the total insert fragments between 260 and 440 bp, and then subtracted the proportion of insert fragments between 326 and 440 bp out of the total insert fragments between 260 and 440 bp, denoted as FS. dist. di. Similarly, significant differences were found between the FS.dist. di of cancer samples and healthy individuals (FIG. 12) . Furthermore, it was observed from FIG. 13 that there may be complementary relationships between FS. dist and FS. dist. di, and a Spearman correlation analysis yielded a correlation coefficient of 0.743 between the two.
[0205] The entire genome was uniformly divided into 100kb-sized regions (bins) and counted the number of reads within each region that fell within a specific short fragment range (e.g., 100-150bp) , referred to as “short fragment count. ” Additionally, they tallied the number of reads within each region falling within a specific long fragment range (e.g., 151-220bp) , termed as “long fragment count. ” Bins were filtered based on GC content or N ratio, among other factors. For example, bin filtering criteria included: 1) mappability >0.6; 2) N ratio <0.5; 3) exclusion of X and Y chromosomes.
[0206] Each bin underwent correction for N ratio, GC content, and mappability. Subsequently, the counts of short and long fragments were individually adjusted using methods such as LOESS to perform the correction.
[0207] In order to reduce the impact of background noise and improve the signal-to-noise ratio, adjacent bins can be combined into larger bins (e.g., 1M, 5M) , and even into the long and short arms of chromosomes (e.g., 1q, 1p, 2q, 2p, 3q, 3p... ) . In this particular implementation, they were merged into 5-Mb bins. Based on the number of short fragments in each 5-Mb bin from normal samples, the 5-Mb bins were filtered, such as the total number of short fragments in one bin was extremely low (Z-score < 3) .
[0208] For each filtered 5-Mb bin, the ratio of short fragments to long fragments was calculated. The absolute sum of the differences between the ratio of each bin and the median value of the ratio of all bins is defined as FS. sdsum. This ratio was also standardized using the mean and standard deviation derived from normal samples in the same bins to compute the Z-score for each 5Mb-bins. The number of 5-Mb bins in each sample with absolute Z-scores greater than 3 was defined as FS. outlier score. As shown in FIG. 14, the differences in FS. outlier score between cancer and normal samples were evident, with a one-tailed t-test p-value close to 0 (P-value < 2.2e-16) , demonstrating an extremely significant difference between the two groups.
[0209] Among the 1197 cancer and normal samples, the performance of above-mentioned FS. dist, FS.dist. di, FS. sdsum and FS. outlier for cancer predicting were shown in FIG. 15.
[0210] To streamline features and prevent overfitting, PCA was applied to the Z-scores of all 5-Mb bins. Ultimately, the top 10 principal components, along with the previously mentioned FS. dist, FS.dist. di, FS. sdsum, and FS. outlier were selected to construct the cancer prediction model using machine learning methods such as SVM, Lasso, GBM, and RF methods. A total of 1197 cases of cancer and normal samples were stratified into two groups based on their peak positions in the fragment size histogram distribution. Separate cancer prediction models were then built for each group using 10-fold cross-validation. This entire process was repeated 3 / 10 / 50 / 100 times, and the average predicted value from these models was designated as the FS value. The performance of each method was evaluated, with the RF method showing an average AUC value of 0.796 (FIG. 16) .
[0211] Example 7: Predict cancer samples using the EMS of the end motif feature
[0212] After obtaining cfDNA sequencing data from plasma samples using high-throughput sequencing (Illumina / MGI) according to Example 1, commonly used software (such as cutadapt and / or SOAPnuke) was employed for sequencing data filtering, alignment to the reference genome of the target species (human hg38 / hg19) , filtering out duplicate sequences, and retaining reads with alignment quality scores greater than or equal to 30, uniquely aligned to the reference genome. Following the example in FIG. 17, a few bases at the 5’ end (1-8bp) or a certain number of bases upstream and downstream of breakpoints (1-6bp upstream and downstream of the breakpoint) were extracted as the end motif. Specifically: as shown in FIG. 17, obtaining information about the 5’ end 2bp bases of the sequencing reads referred to the bases CG; obtaining information about the 5’ end 3bp bases of the sequencing reads referred to the bases CGA; obtaining information about the 5’ end 4bp bases of the sequencing reads referred to the bases CGAC. Obtaining information about the upstream and downstream 2bp bases of the breakpoints of the sequencing reads of the cfDNA referred to the bases CGTC, which meant taking 2bp upstream and downstream, combined for 4bp of base information; obtaining information about the upstream and downstream 3bp bases of the breakpoints of the sequencing reads of the cfDNA referred to the bases ACGTCG, which meant taking 3bp upstream and downstream, combined for 6bp of base information.
[0213] In this example, 5’ end 4 bases were extracted as an end motif (the number of bases could be modified according to requirements) . The length of each cfDNA fragment and its corresponding end motif were obtained. Finally, the number and proportion of the end motif sequences of this group of cfDNA fragments under different fragment lengths were selected.
[0214] The above method was applied to determine the end motif features of all cfDNA fragments from 129 plasma samples of liver cancer patients and 188 samples from healthy individuals. The samples were then divided into training and validation sets, as shown in Table 3:
[0215] Table 3: Sample size of training and validation cohorts
[0216] Obtained the EMS for the degree of deviation in the end distribution of the test plasma sample using the newly developed EMS calculation method.
[0217] (Pi represents the proportion of a specific end motif sequence i in all end motif sequences, while represents the average (or median) proportion of a specific end motif sequence i among all end motif sequences in the normal population. )
[0218] The closer the EMS score is to 1, the more the frequency distribution of different end motifs in the corresponding sample resembles that of the normal population, whereas a score farther from 1 indicates a larger difference, as shown in FIG. 18.
[0219] The EMS scores were evaluated using the ROC curve method to assess the end motif features of the test plasma sample, aiding in determining the source of the test plasma sample. The evaluation set included training data (113 healthy individuals and 78 liver cancer patients) and validation data (75 healthy individuals and 51 liver cancer patients) . Results showed that in the training data, the AUC value for the EMS method was 0.853, while the AUC value for the MDS entropy method (referencing Denis Lo’s article) was 0.738. The Delong test results indicated that the AUC value of the EMS method was significantly higher than that of the MDS method (P-value = 0.003105, Delong test) . In the validation data, the AUC value for the EMS method was 0.812, and the AUC value for the MDS method was 0.708. Delong test results also displayed that the AUC value of the EMS method was significantly higher than that of the MDS method (P-value = 0.001337, Delong test) , as shown in FIG. 19. Taken together, these results indicate that the score for the degree of deviation in the end motif distribution of all cfDNA fragments (EMS) can be applied to cancer screening, with superior screening performance compared to the MDS method.
[0220] Our study (Circulating cell-free DNA fragmentation is a stepwise and conserved process linked to apoptosis. BMC Biol. 2023; 21 (1) : 253) found that apoptosis is a conserved and orderly process, involving different enzymes that cleave DNA at different stages, resulting in DNA fragments of varying lengths with distinct base preferences. For example, the DNASE1L3 cleavage enzyme tends to produce C-terminal sequence fragments; DNASE1 cleavage enzyme tends to produce T-terminal sequence fragments; DFFB cleavage enzyme tends to produce A-terminal sequence fragments. Therefore, by selecting different target fragments for the above analysis, it is possible to differentiate between cancer and normal individuals.
[0221] According to the above method, only the end motif features of cfDNA fragments in the range of the first nucleosome size (50bp-250bp) were calculated. EMS score closer to 1 indicated smaller differences between the test sample and the normal population, whereas the score farther from 1 indicated larger differences. The results revealed that in the training data, the AUC value for the EMS method was 0.851, while the AUC value for the MDS method was 0.748. Delong test results demonstrated that the AUC value of the EMS method was significantly higher than that of the MDS method (P-value =0.0006459, Delong test) . In the validation data, the AUC value for the EMS method was 0.808, and the AUC value for the MDS method was 0.719. Delong test results also showed that the AUC value of the EMS method was significantly higher than that of the MDS method (P-value = 0.001224, Delong test) , as shown in FIG. 20.
[0222] According to the above method, only the end motif features of cfDNA fragments in the range of the second nucleosome size (259bp-439bp) were calculated. The results showed that in the training data, the AUC value for the EMS method was 0.841, while the AUC value for the MDS method was 0.696. Delong test results indicated that the AUC value of the EMS method was significantly higher than that of the MDS method (P-value = 1.189e-05, Delong test) . In the validation data, the AUC value for the EMS method was 0.803, and the AUC value for the MDS method was 0.694. Delong test results also demonstrated that the AUC value of the EMS method was significantly higher than that of the MDS method (P-value = 0.000701, Delong test) , as shown in FIG. 21.
[0223] Based on the above results, it is evident that for comprehensive end motif features of all cfDNA fragments or within a specific range (such as the first nucleosome, the second nucleosome) , the cancer screening performance based on the EMS method is superior to the MDS method.
[0224] Example 8: Using machine learning to predict cancer samples based on the combined end motif features and EMS
[0225] Based on Example 7, obtain the EMS value and the frequency of each end motif of cfDNA fragments within the size range of the first nucleosome (50bp-250bp) . The Wilcoxon rank-sum test (or t-test) was employed to identify statistically significant differences in the frequency of each end motif between cancer patients and healthy individuals. An end motif with adjustment q-value less than 0.001 was selected to build a classifier to differentiate the cancer patients from the healthy subjects. To minimize the issue of overfitting and reduce the number of end motifs, the feature importance of each end motif was evaluated. Subsequently, we arranged the features in descending order of importance and systematically incorporated them one by one until reaching a plateau.
[0226] Based on the EMS of cfDNA fragments within the size range of the first nucleosome and the frequency of selected end motifs, a machine learning method such as logistic regression was used for modeling. During the modeling process, 10-fold cross-validation was employed and repeated 100 times to enhance detection accuracy, and the average value from the 100 models was defined end motif predicting (EMP) value.
[0227] The training and validation set data used for model training and validation were the same as in Example 7, as shown in Table 3. The performance of cancer prediction based on EMP was evaluated using the ROC assessment method, as depicted in FIG. 22. The AUC value for the training set was 0.993, and for the validation set, it was 0.874. These results indicate that the strategy of modeling based on machine learning methods, using selected features combined with EMS values, can be used for predicting the source of cancer samples, with an improvement in performance on the validation set.
[0228] Example 9: Predicting whether the sample originates from a tumor based on the virus abundance
[0229] A process for predicting whether the sample originates from a tumor based on the virus abundance includes the following steps.
[0230] Extract the reads that did not align to the human reference genome from the unaligned BAM file and realign them to the relevant viral sequences using bowtie2. This primarily involves the use of virus sequences associated with cancer types, such as EBV, HPV, HBV, etc. Subsequently, filter out low-quality alignment reads, such as those with a MAPQ value less than 30. After this filtering step, calculate the viral reads fraction (Virus_RF) per megabase of aligned bases using the following formula:
[0231] Based on the Virus_RF values from normal samples, determine a cutoff value at a certain specificity level (using 99%specificity in this Example) . Samples with values exceeding this threshold can be considered at high risk for cancer, indicating that they may originate from cancer. The final prediction results are depicted in FIG. 23, aligning with real-world scenarios.
[0232] Example 10: Calculating the CRS using a Linear regression model
[0233] In Example 10, the PTM (protein-based POC metric) , as well as CNA (CIN score) , FS (FS value) , end motif values (calculated based on Example 7 to obtain EMS score or based on Example 8 to obtain EMP value) , and the sum of Virus_RF for different viruses are used as inputs for multi-cancer screening. In this Example, a linear regression model was used to integrate the five-dimensional features to calculate the CRS, where CRS= a*CNA + b*FS + c*EndMotif + d*Virus + e*PTM (a, b, c, d, and e represent the weight coefficients for CNA, FS, EndMotif, Virus, and PTM) . For 1197 samples (617 cancer samples and 580 healthy samples) , using the maximum AUC value as the evaluation criterion, a grid search method was employed to determine the optimal weight coefficients for the five dimensions, and then the formula was used to calculate the CRS value. The cancer screening performance of CRS was evaluated using the receiver operating characteristic (ROC) curve evaluation method. When the EndMotif value used the EMS score, the AUC value of CRS is 0.89, as shown in FIG. 24; when the EndMotif value used the EMP value, the AUC value of CRS is 0.934, as shown in FIG. 25. The results indicated that the CRS calculated using a linear regression model that integrates the five-dimensional features can be applied to cancer screening, and the performance of CRS was better when the EndMotif dimension used EMP value.
[0234] Example 11: Computing the CRS using a machine-learning model
[0235] In the above example, the PTM (protein-based POC metric) , as well as CNA (CIN score) , FS (FS value) , and end motif values (calculated based on Example 7 to obtain EMS score or based on Example 8 to obtain EMP value) , along with the sum of Virus_RF for different viruses are used as inputs for multi-cancer screening. In this Example, a machine learning model such as logistic regression is used to integrate the five-dimensional features to calculate the CRS. For 1197 samples (617 cancer samples and 580 healthy samples) , a training set and validation set data are randomly constructed in a 9: 1 ratio. The training set data is used for 100 repeated model constructions to determine the optimal model parameters, while the validation set data is used to measure the model’s performance. When the EndMotif value uses EMS values, the average AUC values for the training set and validation set data are 0.893 and 0.845, respectively, as shown in FIG. 26. When the EndMotif value uses EMP value, the average AUC values for the training set and validation set data are 0.938 and 0.94, respectively, as shown in FIG. 27. The results indicate that the CRS calculated using a machine learning model that integrates the five- dimensional features can be applied to multi-cancer screening, and the performance of CRS is better when the EndMotif dimension selects EMP value.
[0236] Example 12: The result of TOO
[0237] True positive cancer samples predicted from Example 10 or Example 11 were used to build a model to predict the cancer tissue of origin (TOO) . Based on the model ensemble approach, only samples from the eight major cancer types (The sample size of uncommon cancer types was too small) were used to build the TOO prediction model by the random forest (RF) using CNA and FS and end motif respectively. In detail, the concentration of all PTMs was used to build the model based on RF and GBM methods. Due to imbalanced sample size for each cancer type, the downsampling method was employed to balance sample size of each cancer type. The regional short-to-long fragment ratio of each 5-Mb bin was used to construct the FS TOO model with five-fold cross-validation and repeat this process 10 times. Regarding CNA, the log R ratio of each 5-Mb bin was normalized by subtracting the mean and dividing by the standard deviation. The normalization z-score of each 5-Mb bin was used as input to build the CNA TOO model with five-fold cross-validation and repeat this process 10 times. The frequency of 256 end motifs was employed to construct the model with five-fold cross-validation and repeat this process 10 times. TOO predicting values from PTMs were integrated with the average predicted value of each cancer type from CNA, FS and end motif model by ensemble modeling method, which aggregated the prediction of each dimensional TOO model and resulted in the final prediction. The organ with the highest traceability score was selected as the most probable source organ, the next highest as the second most probable source site (Top1 and Top2) , and the standardized predicted score was used as the final TOO score of each organ. The final accuracy rate for Top 1 is 71.51%, and for Top 1+Top 2 is 83.43%.
[0238] Example 13
[0239] In the above example, 3 stage IV lymphoma patients were selected. Blood samples were collected blood samples after treatment. Using the method described in the previous Example, the inventors analyzed the changes in CNA before and after treatment. In one patient (ZL0019) , the degree of CNA variation significantly increased after treatment (see FIG. 28) . Conversely, in patients (ZL0020 and ZL0060) where CNAs were initially detected in the pre-treatment samples, CNAs were no longer detected in the plasma after treatment (see FIG. 28) . Additionally, the analysis showed that the CRS calculated for ZL0019 using the method based on Example 10 significantly increased after treatment. However, the CRS values for ZL0020 and ZL0060 decreased to the normal level. Remarkably, after two treatment cycles, all three patients exhibited a Partial Response (PR) according to image evaluations (CT) . However, after four cycles, ZL0019 progressed to Progressive Disease (PD) , while ZL0060 achieved a Complete Response (CR) , and ZL0020 also reached CR after six treatment cycles (FIG. 29) . These findings illustrated the potential of our method in promptly assessing the effectiveness of cancer treatment, as they align with clinical evaluations following treatment.
[0240] OTHER EMBODIMENTS
[0241] It is to be understood that while the invention has been described in conjunction with the detailed description thereof, the foregoing description is intended to illustrate and not limit the scope of the invention, which is defined by the scope of the appended claims. Other aspects, advantages, and modifications are within the scope of the following claims.
Claims
1.A method for predicting a cancer risk for a subject based on a test sample from the subject, the method comprising:obtaining first measurement data characterizing levels of a set of protein tumor markers in the test sample;determining, using a computer system, a protein-based probability of cancer (POC) metric based at least on the first measurement data;obtaining second measurement data comprising cell-free DNA (cfDNA) sequencing data in the test sample;determining an end motif metric characterizing end motifs of cfDNA fragments in the test sample using at least an end motif score (EMS) characterizing a degree of deviation of a distribution of the end motifs of the test sample from a control distribution;determining a chromosomal instability index characterizing a copy number aberration (CNA) of the cfDNA fragments in the test sample; anddetermining a fragment size (FS) metric characterizing a fragment size pattern of the cfDNA fragments in the test sample;processing a first model input comprising (i) the protein-based POC metric, (ii) the end motif metric, (iii) the chromosomal instability index, and (iv) the FS metric using a first machine learning model to generate prediction data comprising a cancer risk score that characterizes the cancer risk for the subject; andoutputting the prediction data comprising the cancer risk score for the subject.2.The method of claim 1, wherein:the method further comprises determining a virus metric characterizing abundances of one or more cancer related viruses in the test sample; andthe first model input further comprises the virus metric.3.The method of claim 2, wherein the one or more viruses comprise one or more of: HBV, HPV, EBV, HHV, HIV, or JCPyV.4.The method of claim 2 or claim 3, wherein determining the virus metric comprises:obtaining sequencing data characterizing sequences of the cfDNA fragments of the test sample;identifying, from the sequencing data, a first set of sequences that do not align with a human reference genome;identifying, from the first set of sequences, a second set of sequences that aligned with one of the viruses; anddetermining the virus metric based on a ratio of sequencing reads in the second set of sequences out of total sequencing reads.5.The method of any of any preceding claim, wherein determining the protein-based POC metric comprises:processing a second model input comprising the measured levels of the protein tumor markers using a second machine learning model to generate an output that specifies the tumor metric.6.The method of claim 5, further comprising:obtaining one or more clinical parameters of the subject; whereinthe second model input further comprises the one or more clinical parameters of the subject.7.The method of any preceding claim, wherein the set of tumor markers comprises or consists of AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, and CYFRA 21-1.8.The method of any preceding claim, wherein the subject is a male and the set of tumor markers comprises or consists of AFP, CA125, CA15-3, CA19-9, CEA, CYFRA 21-1, ProGRP, SCCA, and PSA; or wherein the subject is a female and the set of tumor markers comprises or consists of AFP, CA125, CA15-3, CA19-9, CEA, CYFRA 21-1, ProGRP, and SCCA.9.The method of claim 5, wherein the second machine learning model combines two or more models selected from the group consisting of generalized linear models (GLMs) , gradient boosting machines (GBMs) , random forests (RFs) , and support vector machines (SVMs) .10.The method of claim 9, wherein the second machine learning model combines the two or more selected models using a linear combination.11.The method of any of claims 5-10, wherein determining the protein-based POC metric further comprises: performing an outlier analysis, wherein a higher level of a protein tumor marker increases the tumor metric.12.The method of any preceding claim, wherein determining the end motif metric comprises:obtaining cfDNA sequencing data characterizing a set of cfDNA sequences of the cfDNA fragments of the test sample;for each of the set of cfDNA sequences, determining a respective length of the cfDNA sequence;for each of the set of cfDNA sequences, extracting a respective end motif as (i) a sub-sequence comprising a first predefined number of bases at a 5' end terminus of the respective cfDNA sequence or (ii) a sub-sequence comprising a second predefined number of bases before or after a breakpoint of the respective cfDNA sequence; andfor each of a set of different cfDNA fragment size, determining a respective distribution of different types of end motifs in the test sample.13.The method of claim 12, wherein the first predefined number of bases is 1bp, or 2bp, or 3bp, or 4bp, or 5bp, or 6bp, or 7bp, or 8bp.14.The method of claim 12 or claim 13, wherein the second predefined number of bases is 1bp, or 2bp, or 3bp, or 4bp, or 5bp, or 6bp.15.The method of any of claims 12-14, wherein determining the end motif metric further comprises:computing the EMS of the test sample based on (i) the distributions of the different types of end motifs for cfDNA fragments of a predefined fragment size range in the test sample and (ii) distributions of different types of end motifs for the predefined fragment size range in control samples.16.The method of claim 15, wherein the predefined fragment size range comprises (i) a first length range of 50bp-250bp; or (ii) a sub-range within the first length range of 50bp-250bp.17.The method of claim 15 or 16, wherein the predefined fragment length range comprises (i) a second length range of 250bp-450bp, or (ii) a sub-range within a second length range of 250bp-450bp.18.The method of any of claims 12-17, wherein the EMS of the test sample is computed as: wherein i represents a type of end motif, Pi a proportion of the end motif i in all end motifs of the cfDNA fragments of the test sample, is an average or median value of the proportion of end motif iin the control samples, m is the total number of end motifs.19.The method of any of claims 12-18, wherein the EMS of the test sample is computed as a weighted sum of (i) a first EMS determined based on a first predefined fragment length range and (ii) a second EMS determined based on a second predefined fragment length range.20.The method of claim 19, wherein weight coefficients of the weighted sum are determined using a third machine learning model.21.The method of any preceding claim, wherein determining the end motif metric comprises:processing a model input comprising (i) the EMS and (ii) occurrence frequencies of a selected subset of end motifs in the test sample using a fourth machine learning model to generate an output comprising the end motif metric.22.The method of claim 21, wherein the selected subset of end motifs is determined using one or more statistical tests that identify differences in end motif occurrence frequencies between healthy subjects and cancer patients.23.The method of any preceding claim, wherein determining the FS metric comprises:obtaining sequencing data characterizing a set of sequences of the cfDNA fragments of the test sample;determining, from the sequencing data, one or more first distribution parameters characterizing proportions of fragments within specific length ranges relative to other length ranges;determining, from the sequencing data, one or more second distribution parameters by performing binning, filtering, combining, and normalization on the sequencing data; andprocessing a fifth model input comprising the first distribution parameters and the second distribution parameters using a fifth machine learning model to generate an output specifying the FS metric.24.The method of any preceding claim, further comprising:before processing the first model input using the first machine learning model, performing a training process to train the first machine learning model using a plurality of training examples, each respective training example comprising (i) a respective training input comprising a respective set of metrics determined based on measurement data from a respective subject and (ii) a respective training output characterizing a diagnosis of the respective subject, the plurality of examples comprising (i) a set of training examples using measurement data from cancer patients and (ii) a set of training examples using measurement data from control subjects.25.A method for predicting a cancer risk for a subject based on a test sample from the subject, the method comprising:obtaining measurement data comprising cell-free DNA (cfDNA) sequencing data in the test sample;determining an end motif metric that comprises an end motif score (EMS) characterizing a degree of deviation of a distribution of end motifs of cfDNA fragments in the test sample from a control distribution;determining a cancer risk score that characterizes the cancer risk for the subject from at least the EMS; andoutputting the prediction data comprising the cancer risk score for the subject.26.The method of claim 25, wherein the EMS is determined according to any of claims 12-20.27.The method of claim 25 or 26, wherein determining the cancer risk score comprises:processing a model input comprising (i) the EMS and (ii) occurrence frequencies of a selected subset of end motifs in the test sample using a fourth machine learning model to generate an output comprising the cancer risk score.28.The method of 27, wherein the selected subset of end motifs is determined using one or more statistical tests that identify differences in end motif occurrence frequencies between healthy subjects and cancer patients.29.The method of any preceding claim, further comprising:in response to the cancer risk score exceeding a predefined threshold, predicting a tissue of origin for a cancer patient using a sixth machine learning model.30.The method of any preceding claim, further comprising:using the prediction data to monitor a cancer treatment of the subject.31.A system comprising:one or more computers; andone or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations of the respective method of any one of claims 1-30.32.One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations of the respective method of any one of claims 1-30.