Systems and methods for detection of non-coding rnas
The use of non-coding RNA molecules in liquid biopsies, analyzed by a trained algorithm, addresses the limitations of current MRD detection methods by enhancing sensitivity and enabling timely monitoring and prediction of cancer recurrence without tumor tissue sequencing.
Patent Information
- Application Number
- PCT/US2025/023242
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-04
- Filing Date
- 2025-04-04
- Publication Date
- 2025-10-09
AI Technical Summary
Current MRD detection methods for cancers like TNBC are limited by the need for tumor tissue sequencing, which is challenging to obtain and delays therapy, and cell-free RNA modalities lack sensitivity, especially for early-stage tumors.
A method involving the use of non-coding RNA molecules, specifically orphan non-coding RNAs (oncRNAs), which are sequenced and analyzed using a trained algorithm to detect residual cancer cells in liquid biopsies, providing a classification report.
Enhances sensitivity for MRD detection, allowing for early-stage tumor monitoring and predicting clinical characteristics and recurrence risk without the need for tumor tissue sequencing, thereby facilitating timely treatment decisions.
Smart Images

Figure US2025023242_09102025_PF_FP_ABST
Abstract
Description
SYSTEMS AND METHODS FOR DETECTION OF NON-CODING RNASCROSS REFERENCE
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 574,826, filed April 4, 2024, which is incorporated herein by reference for all purposes.BACKGROUND
[0002] Residual disease, also known as minimal residual disease (MRD) after neoadjuvant chemotherapy (NACT) is associated with a high risk of recurrence in triplenegative breast cancer (TNBC). Early detection and treatment of MRD in TNBC can significantly increase survival.SUMMARY
[0003] Cell-free circulating tumor DNA (ctDNA) assays are used for DNA profiling of tumor DNA variants for patient treatment selection and pharmacodynamics for targeted therapies. Baseline patient tumor burden, estimated using genotyping, epigenomic or fragmentomic methods has been shown to be negatively associated with patient prognosis and overall survival. The change in tumor burden between pre-treatment and on-treatment timepoints, known as molecular response, has been shown to predict response to treatment and overall survival, and is being investigated as a surrogate endpoint in clinical trials.
[0004] CtDNA assays are also used following curative intent therapy to inform presence of minimal residual disease (MRD) in the neoadjuvant or adjuvant setting, for eventual escalation or de-escalation of treatment. The majority of MRD detection assays require sequencing the tumor tissue to identify tumor DNA variants to track in blood (tumor-informed). However, tumor tissue is challenging to obtain for many patients (e.g. lung cancer patients) and adds several weeks of processing time, which delays therapy and is not practical for clinical management. CtDNA assays that do not require sequencing the tumor tissue first (tumor-naive) struggle with sensitivity, which is dependent on the level of tumor shedding into the blood, particularly in early stage tumors, where tumor signal is low.
[0005] Furthermore, clinical characteristics of tumors are often categorical variables (grade, stage, and TNM scores). Continuous measures of tumors, such as tumor size, do not reflect tumor invasiveness and phenotypic characteristics. Large tumors, for example, can originate from both early stage (stage I) or late stage (stage IV) lesions. Developing machine learning models for a better understanding of tumor burden, therefore, requires a continuous measurement reflecting commonly available clinical characteristics.
[0006] There is a need for more sensitive analytes and modalities to detect and monitor cancer particularly in patients in the early stage setting. Currently, commercially available liquid biopsy tests rely only on circulating tumor DNA. Cell-free RNA modalities have not demonstrated success thus far as a tool for monitoring and MRD detection in oncology.
[0007] In an aspect, disclosed herein is a method, comprising: (a) providing a sample from a subject, wherein the subject has received treatment for cancer and has or is suspected of having residual cancer cells of the cancer, and wherein the sample comprises one or more noncoding RNA molecules that are present in the residual cancer cells and have a 90th percentile expression in non-cancerous cells of below 0.5 count per-million reads (cpm), (b) sequencing the one or more non-coding RNA molecules, thereby generating a data set comprising data corresponding to a presence of the one or more non-coding RNA molecules in the sample, (c) inputting the data set into a trained algorithm to generate a classification that includes an indication that the sample is positive or negative for the residual cancer cells, and (d) electronically outputting the classification in a report.
[0008] In some embodiments, the method further comprises (e) providing the report to a user. In some embodiments, the user is a clinician. In some embodiments, the providing further comprises delivering the report to a computer interface of the user.
[0009] In another aspect, disclosed herein is a method, comprising: (a) providing a sample from a subject, wherein the subject has received treatment for cancer and has or is suspected of having residual cancer cells of the cancer, and wherein the sample comprises one or more non-coding RNA molecules having a sequence of any one of SEQ ID NOs: 1-10,334, (b) sequencing the one or more non-coding RNA molecules, thereby generating a data set comprising data corresponding to a presence of the one or more non-coding RNA molecules in the sample, (c) inputting the data set into a trained algorithm to generate a classification that includes an indication that the sample is positive or negative for the residual cancer cells, and (d) electronically outputting the classification in a report.
[0010] In some embodiments, the method further comprises (e) providing the report to a user. In some embodiments, the user is a clinician. In some embodiments, the providing further comprises delivering the report to a computer interface of the user.
[0011] In another aspect, disclosed herein is a method, comprising: (a) providing a sample from a subject, wherein the subject has received treatment for cancer and has or is suspected of having residual cancer cells of the cancer, wherein the sample comprises one ormore non-coding RNA molecules, (b) sequencing the sample, thereby obtaining a first data set comprising a first set of one or more non-coding RNA sequences of the one or more non-coding RNA molecules in the sample, (c) providing a second data set comprising a second set of one or more non-coding RNA sequences of one or more non-coding RNA molecules that are (i) present in residual cancer cells and (ii) have a 90th percentile expression in normal samples below 0.5 count-per-million (cpm) reads, (d) providing the first data set and the second data set into a trained algorithm to generate a classification that includes an indication that the sample is positive or negative for the residual cancer cells, and (e) electronically outputting the classification in a report.
[0012] In some embodiments, the method further comprises (f) providing the report to a user. In some embodiments, the user is a clinician. In some embodiments, the providing further comprises delivering the report to a computer interface of the user.
[0013] In another aspect, disclosed herein is a method, comprising: (a) providing a sample from a subject, wherein the subject has received treatment for cancer and has or is suspected of having residual cancer cells of the cancer, wherein the sample comprises one or more non-coding RNA molecules, (b) sequencing the sample, thereby obtaining a first data set comprising a first set of one or more non-coding RNA sequences of the one or more non-coding RNA molecules in the sample, (c) providing a second data set comprising a second set of one or more non-coding RNA sequences of one or more non-coding RNA molecules having a sequence of any one of SEQ ID NOs: 1-10,334, (d) providing the first data set and the second data set into a trained algorithm to generate a classification that includes an indication that the sample is positive or negative for the residual cancer cells, and (e) electronically outputting the classification in a report.
[0014] In some embodiments, the method further comprises (f) providing the report to a user. In some embodiments, the user is a clinician. In some embodiments, the providing further comprises delivering the report to a computer interface of the user.
[0015] In some aspects, the classification further comprises a prediction of a likelihood that the residual cancer cells will proliferate. In some aspects, the classification further comprises a prediction of survival of the subject. In some aspects, the classification further comprises a prediction of a clinical characteristic comprising a cancer stage, a tumor (T) stage, a node invasion score, a metastasis score, a tumor grade score, a tumor size score, a cancer site score, or any combination thereof. In some aspects, the classification further comprises a prediction of a pathologic complete response (pCR).
[0016] In some aspects, the treatment comprises a neoadjuvant chemotherapy.
[0017] In some aspects, the one or more non-coding RNA molecules have a length of less than 200 nucleotides. In some aspects, the one or more non-coding RNA molecules have a length of between 50 and 100 nucleotides.
[0018] In some aspects, the sequencing in (b) comprises subjecting the one or more non-coding RNA molecules to reverse transcription to generate one or more complementary deoxyribonucleic acid (DNA) molecules. In some aspects, the sequencing in (b) further comprises sequencing the one or more cDNA molecules or derivatives thereof. In some aspects, the sequencing in (b) further comprises sequencing-by-synthesis.
[0019] In some aspects, the methods further comprise, after (b), amplifying a cDNA molecule of the one or more cDNA molecules. In some aspects, the amplifying comprises performing a polymerase chain reaction (PCR). In some aspects, the amplifying comprises rolling circle amplification.
[0020] In some aspects, the sample is a cell-free sample. In some aspects, the sample comprises serum. In some aspects, the sample comprises plasma. In some aspects, the sample comprises urine. In some aspects, the sample comprises lymph. In some aspects, the sample comprises saliva. In some aspects, a total volume of the sample is between 20 microliters and 2 milliliters. In some aspects, the sample comprises plasma, and wherein a total volume of the sample is between 100 microliters and 1 milliliter.
[0021] In some aspects, the trained algorithm comprises a machine learning model. In some aspects, the machine learning model comprises a neural network. In some aspects, the neural network comprises a regressor neural network, a multi-head neural network, or any combination thereof.
[0022] In some aspects, the trained algorithm is trained using non-coding RNA derived from a treatment-naive patient sample. In some aspects, the treatment-naive patient sample comprises serum. In some aspects, the trained algorithm is trained across multiple demographics, cancer types, and cancer stages. In some aspects, the trained learning algorithm is trained on clinical characteristics comprising a cancer stage, a T stage, a node invasion score, a metastasis score, a tumor grade score, a tumor size score, a cancer site score, or any combination thereof.
[0023] In some aspects, the residual cancer cells comprise breast cancer cells. In some aspects, the breast cancer cells comprise triple-negative breast cancer cells.INCORPORATION BY REFERENCE
[0024] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also “figure” and “FIG.” herein), of which:
[0026] FIG. 1 shows an example of a method of using a trained machine learning algorithm (or generative Al model), as described herein, to perform a multi-class prediction of a metastatic status, a nodal status, a tumor grade, a tumor burden status
[0027] FIG. 2 shows a method of training a machine learning algorithm, as described herein, to predict a tumor burden from oncRNAs derived from a cell-free patient sample.
[0028] FIG. 3A illustrates systems and methods described herein to indicate cancer presence, residual cancer cell presence, or tumor absence in a liquid biopsy sample obtained from a patient.
[0029] FIG. 3B shows the performance of a trained machine learning algorithm to predict tumor size.
[0030] FIGs. 4A-4C show the performance of a trained machine learning algorithm to predict tumor burden and tumor grade. FIG. 4A shows predicted tumor size plotted against actual tumor grade (T score). FIG. 4B shows actual tumor size plotted against actual tumor grade (T score). FIG. 4C shows a correlation between predicted tumor burden and tumor stage, with predicted tumor size plotted against actual tumor stage.
[0031] FIG. 5 shows a concordance between a metastatic status predicted by a trained machine learning algorithm and actual metastatic status.
[0032] FIG. 6 shows a concordance index (y-axis) as a measure of the performance of a multi-variate Cox proportional hazards model in predicting patient overall survival from a tumor burden output.
[0033] FIG. 7A shows a significant, positive correlation of predicted tumor burden (x-axis) with actual tumor size (y-axis) as generated by a trained machine learning algorithm, as described herein.
[0034] FIG. 7B shows a significant, positive correlation of predicted tumor burden (x-axis) with computational tumor burden (y-axis) as generated by a trained machine learning algorithm, as described herein..
[0035] FIG. 7C shows the area under the receiver-operating characteristic curve for distinguishing cancer oncRNA samples from control oncRNA samples, with a sensitivity of 64.1% at 90% specificity using a trained machine learning algorithm, as described herein.
[0036] FIG. 7D shows the sensitivity for distinguishing tumors of different stages at 90% specificity using the trained machine learning algorithm. FIG. 7E shows the sensitivity for distinguishing tumors of different T score at 90% specificity using a trained machine learning algorithm as described herein.
[0037] FIG. 8 shows cohort demographics, BRCA1 / 2 mutation status, T stage, nodal status, TNM stage, RCB class, adjuvant radiation status, and adjuvant chemotherapy status in a study investigate the capability of an oncRNA signature to predict a risk of recurrence of residual disease (RD) in triple-negative breast cancer (TNBC) patients with RD
[0038] FIG. 9 shows an association of overall survival and event-free survival with oncRNA risk status by univariable analysis.
[0039] FIG. 10A shows an association of overall survival with oncRNA risk status.
[0040] FIG. 10B shows an association of event-free survival with oncRNA risk status.
[0041] FIG. 10C shows oncRNA low-risk status was associated with better event- free survival outcomes.
[0042] FIG. 10D shows oncRNA low-risk status was associated with better overall survival.
[0043] FIG. 11 shows the design of a study for predicting distant recurrence-free survival (DRFS) in breast cancer patients using a standard blood draw
[0044] FIG. 12 shows the demographics of a study for predicting distant recurrence- free survival (DRFS) in breast cancer patients using a standard blood draw.
[0045] FIG. 13 shows a breakdown of patient receptor subtypes by treatment arm for a study for predicting distant recurrence-free survival (DRFS) in breast cancer patients using a standard blood draw.
[0046] FIG. 14 shows data that indicate continuous oncRNA risk scores are predictive of distant recurrence in univariate and multivariate models.
[0047] FIGs. 15A-15D show Kaplan-Meier survival curves for predicting DRFS when stratifying by pCR alone (FIG. 15A), oncRNA risk alone (FIG. 15B), or pCR further stratified by oncRNA risk score group within patients that achieved pCR (FIG. 15C), and patients that did not achieve pCR (FIG. 15D).
[0048] FIG. 16 shows a computer system that is programmed or otherwise configured to implement methods provided herein.DETAILED DESCRIPTION
[0049] While various embodiments of the invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed.
[0050] The present disclosure provides systems and methods for analyzing small RNA sequences that serve as biomarkers which are indicative of residual cancer cells in a patient who has received treatment for a cancer. In some embodiments, the methods comprise use of cell-free small RNAs in a suitable sample. In some embodiments, the sample is a biological sample. In some embodiments, the biological sample is a whole blood, serum, or plasma sample.
[0051] Orphan non-coding RNAs (oncRNAs) can be found in residual cancer tissue and cancer tissue in higher levels than baseline, but are typically not found in noncancerous tissue at high levels. Specific oncRNAs can be present in different amounts in different types of cancers. Detection of a combination of two or more of oncRNAs can result in a “signature” of a unique oncRNA combination that can be indicative of a cancer, a pre-cancerous condition, or residual cancer cells remaining in a patient after a cancer treatment (e.g., MRD). oncRNAs can be used in a liquid biopsy platform for sensitive and accurate early detection of MRD. oncRNAs can be used to detect a cancer, and can predict clinico-pathological characteristics of the cancer. The systems and methods described herein can be used to monitor patients after receiving a cancer treatment to detect a presence or an absence of MRD. In some embodiments, the systems and methods described herein can be used to determine a risk of recurrence of cancer in a patient having, or is suspected of having, MRD.
[0052] Methods and systems of the present disclosure may employ, unless otherwise indicated, techniques of molecular biology (including recombinant techniques), microbiology, cell biology, and biochemistry. Such techniques may be described by, for example, “Molecular Cloning: A Laboratory Manual”, 2ndedition (Sambrook et al., 1989); “Oligonucleotide Synthesis” (M.J. Gait, ed., 1984); “Animal Cell Culture” (R.I. Freshney, ed., 1987); “Methods in Enzymology” (Academic Press, Inc.); “Handbook of Experimental Immunology”, 4thedition (D.M. Weir & C.C. Blackwell, eds., Blackwell Science Inc., 1987); “Gene Transfer Vectors for Mammalian Cells” (J.M. Miller & M.P. Calos, eds., 1987); “Current Protocols in Molecular Biology” (F.M. Ausubel et al., eds., 1987); and “PCR: The Polymerase Chain Reaction”, (Mullis et al., eds., 1994), each of which is incorporated by reference in its entirety.
[0053] Sequencing methods can be used to generate sequencing reads derived small RNA. For example, small RNAs can be detected using Next Generation Sequencing (NGS), e.g., Sequencing-By-Synthesis using, for example, the NovaSeq, Next Seq, HiSeq, HiScan, GenomeAnalyzer, or MiSeq systems (Illumina, Inc., San Diego CA) or the Ultima Genomics Sequencer (Ultima Genomics, Newark CA). Small RNAs can also be detected using Ion Torrent Sequencing (Ion Torrent Systems, Inc., Gulliford, Conn.) or other suitable methods of semiconductor sequencing.
[0054] Determining a level of small RNA biomarkers may comprise counting a number of sequencing reads comprising a small RNA biomarker. The number of sequencing reads that comprise a given biomarker may correspond to an amount of a small RNA biomarker in a sample.
[0055] As used in the present disclosure and claims, the singular forms “a”, “an” and “the” include plural forms unless the context clearly dictates otherwise.
[0056] Whenever the term “at least,” “greater than,” or “greater than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “at least” or “greater than” applies to each one of the numerical values in that series of numerical values.
[0057] Whenever the term “no more than,” “less than,” or “less than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “no more than” or “less than” applies to each one of the numerical values in that series of numerical values.
[0058] It is understood that wherever embodiments are described herein with the language “comprising” otherwise analogous embodiments described in terms of “consisting of’ and / or “consisting essentially of’ are also provided. It is also understood that wherever embodiments are described herein with the language “consisting essentially of’ otherwise analogous embodiments described in terms of “consisting of’ are also provided.
[0059] The term “and / or” as used in a phrase such as “A and / or B” herein is intended to include both A and B; A or B; A (alone); and B (alone). Likewise, the term “and / or” as used in a phrase such as “A, B, and / or C” is intended to encompass each of the following embodiments: A, B, and C; A, B, or C; A or C; A or B; B or C; A and C; A and B; B and C; A (alone); B (alone); and C (alone).
[0060] The term “about” or “approximately” as used herein is meant to refer to within 5%, within 4%, within 3%, within 2%, within 1%, of a given value or range.
[0061] The term “biomarker”, as used herein, generally refers to a biological molecule present in an individual at varying concentrations useful in predicting the cancer status of an individual. A biomarker may include but is not limited to, nucleic acids, proteins and variants and fragments thereof. A biomarker may be DNA comprising the entire or partial nucleic acid sequence encoding the biomarker, or the complement of such a sequence. Biomarker nucleic acids useful in the systems and methods of the present disclosure are considered to include both DNA and RNA comprising the entire or partial sequence of any of the nucleic acid sequences of interest.
[0062] The term “bodily fluid”, as used herein, generally refers to a biological sample in the form of a bodily fluid comprising RNA (RNA) including blood (or a fraction ofblood such as plasma or serum), lymph, mucus, tears, saliva, sputum, urine, semen, stool, CSF (cerebrospinal fluid), breast milk, and ascites fluid. In some embodiments, the bodily fluid is urine. In some embodiments, the bodily fluid is fractionated serum comprising extracellular vesicles such as exosomes.
[0063] The terms “cancer” and “cancerous”, as used herein, generally refer to or describe the physiological condition in mammals in which a population of cells are characterized by unregulated cell growth. In some embodiments, the cancer is a breast cancer. In some embodiments, the cancer is a liver cancer. In some embodiments, the cancer is a gastric cancer. In some embodiments, the cancer is a lung cancer. In some embodiments, the cancer is a colorectal cancer. In some embodiments, the cancer is a kidney cancer.
[0064] The terms “minimal residual disease” and “residual cancer”, as used herein, generally refer to or describe the physiological condition in mammals in which one or more cancer cells or one or more cancer tissues are present in the mammal after the mammal has received a treatment for the cancer. In some embodiments, the residual cancer is a breast cancer. In some embodiments, the residual cancer is a liver cancer. In some embodiments, the residual cancer is a gastric cancer. In some embodiments, the residual cancer is a lung cancer. In some embodiments, the residual cancer is a colorectal cancer. In some embodiments, the residual cancer is a kidney cancer. In some embodiments, the breast cancer is triple-negative breast cancer (TNBC). In some embodiments, the treatment is a chemotherapy. In some embodiments, the treatment is a radiation therapy. In some embodiments, the treatment is an immunotherapy. In some embodiments, the treatment is a hormone therapy. In some embodiments, the treatment is a stem cell therapy. In some embodiments, the treatment is a gene therapy. In some embodiments, the chemotherapy is a neoadjuvant chemotherapy (NACT).
[0065] The term "correlate" or "correlating", as used herein, generally refers to a statistical association between instances of two events, where events may include numbers, data sets, and the like. For example, when the events involve numbers, a positive correlation (also referred to herein as a "direct correlation") means that as one increases, the other increases as well. A negative correlation (also referred to herein as an "inverse correlation") means that as one increases, the other decreases. The present disclosure provides small RNAs, the levels of which are correlated with a particular outcome measure, such as between the level of a small RNA and the likelihood of recurrence in MRD. For example, the increased level of a small RNA may be negatively correlated with a likelihood of good clinical outcome for the patient. In this case, for example, the patient may have a decreased likelihood of long-term survival withoutrecurrence of the cancer and / or a positive response to a chemotherapy, and the like. Such a negative correlation indicates that the patient likely has a poor prognosis or may respond poorly to a chemotherapy, and this may be demonstrated statistically in various ways, e.g., by a high hazard ratio.
[0066] The terms “identical” or “percent identity” or “homology” in the context of two or more nucleic acids, as used herein, generally refer to two or more sequences or subsequences that are the same or have a specified percentage of nucleotides or amino acid residues that are the same, when compared and aligned (introducing gaps, if necessary) for maximum correspondence, not considering any conservative amino acid substitutions as part of the sequence identity. The percent identity may be measured using sequence comparison software or algorithms or by visual inspection. Various algorithms and software that may be used to obtain alignments of amino acid or nucleotide sequences. These include, but are not limited to, BLAST, ALIGN, Megalign, BestFit, GCG Wisconsin Package, and variations thereof. In some embodiments, two nucleic acids of the present disclosure are substantially identical, meaning they have at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, and in some embodiments at least about 95%, 96%, 97%, 98%, 99% nucleotide or amino acid residue sequence identity, when compared and aligned for maximum correspondence, as measured using a sequence comparison algorithm or by visual inspection. In some embodiments, identity exists over a region of the sequences that is at least about 10, at least about 20, at least about 40-60 nucleotides, at least about 60-80 nucleotides or any integral value therebetween. In some embodiments, identity exists over a longer region than 60-80 nucleotides, such as at least about 80-100 nucleotides, and in some embodiments the sequences are substantially identical over the full length of the sequences being compared.
[0067] The term “level”, as used herein, generally refers to qualitative or quantitative determination of the number of copies of a non-coding RNA transcript. An RNA transcript exhibits an “increased level” when the level of the RNA transcript is higher in a first sample, such as in a clinically relevant subpopulation of patients (e.g., patients who have cancer), than in a second sample, such as in a related subpopulation (e.g., patients who do not have cancer). In the context of an analysis of a level of an RNA transcript in a tumor sample obtained from an individual patient, an RNA transcript exhibits “increased level” when the level of the RNA transcript in the subject trends toward, or more closely approximates, the level characteristic of a clinically relevant subpopulation of patients. The methods herein are useful in determining levels of certain small ncRNA sequences as well as the presence or absence of selected small ncRNA biomarkers and the absolute number of distinct small ncRNA sequences.
[0068] The term “metastasis”, as used herein, generally refers to the process by which a cancer spreads or transfers from the site of origin to other regions of the body with the development of a similar cancerous lesion at a new location. A “metastatic” or“ metastasizing” cell is one that loses adhesive contacts with neighboring cells and migrates (e.g., via the bloodstream or lymph) from the primary site of disease to secondary sites.
[0069] As generally used herein, the term “orphan non-coding RNA” or “oncRNA” refers to a category of small RNAs (smRNAs), typically small non-coding RNAs (small ncRNAs) that are present in tumors but largely absent in healthy tissue. By way of non-limiting example only, oncRNAs may refer to small ncRNAs that (i) have a CPM below 0.09 in 95% of normal serum samples; (ii) have an adjusted p-value of < 0.1 following an association study of tumor tissue versus normal tissues using a generalized linear model and correcting for known confounders, including age and sex; and (iii) as for small RNAs generally, are less than about 200 nt in length, such as in the range of 50-100 nt. In some cases, the term oncRNAs have a 90th percentile expression in non-cancerous (or non-TNBC) cells of below 0.5 count per-million cells. In addition, while an oncRNA species is a non-coding RNA sequence, it may overlap, in part, with an adjacent coding sequence. Representative embodiments of detection and / or quantification of oncRNA molecules in a sample of a subject may be found in PCT International Application Pub. WO 2019 / 094781 and PCT International Application Pub. WO 2022 / 040106, both of which are expressly incorporated by reference herein in entirety.
[0070] A “patient response” , as used herein, generally may be assessed using any endpoint indicating a benefit to the patient, including, without limitation, (1) inhibition, to some extent, of tumor growth, including slowing down and complete growth arrest; (2) reduction in the number of tumor cells; (3) reduction in tumor size; (4) inhibition (i.e., reduction, slowing down or complete stopping) of tumor cell infiltration into adjacent peripheral organs and / or tissues; (5) inhibition (i.e. reduction, slowing down or complete stopping) of metastasis; (6) enhancement of anti-tumor immune response, which may, but does not have to, result in the regression or rejection of the tumor; (7) relief, to some extent, of one or more symptoms associated with the cancer; (8) increase in the length of survival following treatment; (9) reduction in tumor recurrence after a cancer treatment; and / or (10) decreased mortality at a given point of time following treatment.
[0071] The terms “polynucleotide” and “nucleic acid” and “nucleic acid molecule”, as used herein, generally refer to polymers of nucleotides of any length, and include DNA and RNA. The polynucleotides can be deoxyribonucleotides, ribonucleotides, modified nucleotidesor bases, and / or their analogs, or any substrate that can be incorporated into a polymer by DNA or RNA polymerase.
[0072] The terms “polypeptide” and “peptide” and “protein”, as used herein, generally refer to polymers of amino acids of any length. The polymer may be linear or branched, it may comprise modified amino acids, and it may be interrupted by non-amino acids. The terms also encompass an amino acid polymer that has been modified naturally or by intervention; for example, disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation or modification, such as conjugation with a labeling component. Also included within the definition are, for example, polypeptides containing one or more analogs of an amino acid (including, for example, unnatural amino acids), as well as other modifications. It is understood that, because the polypeptides of the present disclosure may be based upon antibodies or fusion proteins, in certain embodiments, the polypeptides can occur as single chains or associated chains (e.g., dimers).
[0073] The term "prognosis", as used herein, generally refers to the prediction of the likelihood of cancer-attributable death or progression, including recurrence, metastatic spread, and drug resistance, of neoplastic disease, such as breast cancer.
[0074] The term “reference” RNA transcript, as used herein, generally refers to an RNA transcript whose level can be used to compare the level of an RNA transcript in a test sample. In an embodiment, reference RNA transcripts include housekeeping genes, such as betaglobin, alcohol dehydrogenase, or any other RNA transcript, the level or expression of which does not vary depending on the disease status of the cell containing the RNA transcript. In another embodiment, all of the assayed RNA transcripts, or a subset thereof, may serve as reference RNA transcripts.
[0075] The term “small RNA” (smRNA), as used herein, generally refers to RNA that has a length of less than 200 nucleotides and includes transfer RNA (tRNA), ribosomal RNA (rRNA), snoRNAs, microRNA (miRNA), siRNAs, small nuclear (snRNA), Y RNA, vault RNA, antisense RNA, tiRNA (transcription initiation RNA), TSSa-RNA (transcriptional startsite associated RNA) and piwiRNA (piRNA). It is to be understood that a "non-coding RNA" may overlap, in part, with a segment of RNA that does code for a protein, although the entirety of the non-coding RNA sequence per se does not. In some embodiments, a small smRNA as used herein is between 50 and 100 nucleotides. A smRNA may be of endogenous origin (e.g., a human small non-coding RNA) or exogenous origin (e.g., virus, bacteria, parasite). In someembodiments, any of the methods disclosed herein comprise detecting any one or combination of RNAs disclosed above.
[0076] The term “subject”, as used herein, generally refers to any animal (e.g., a mammal), including, but not limited to, humans, non-human primates, canines, felines, rodents, and the like. In some embodiments, the subject is a human subject. The terms "subject," "individual," and "patient" are used interchangeably herein. The terms "subject," "individual," and "patient" thus encompass individuals having cancer (e.g., breast cancer), including those who have undergone or are candidates for a cancer treatment to treat or remove cancerous tissue.
[0077] The term “therapeutically effective amount”, as used herein, generally refers to a quantity sufficient to achieve a desired therapeutic effect, for example, an amount which results in the prevention or amelioration of or a decrease in the symptoms associated with a disease that is being treated, e.g., disorders associated with cancer growth or a hyperproliferative disorder. The amount of compound administered to the subject will depend on the type and severity of the disease and on the characteristics of the individual, such as general health, age, sex, body weight and tolerance to drugs. It will also depend on the degree, severity and type of disease. The skilled artisan will be able to determine appropriate dosages depending on these and other factors. The regimen of administration can affect what constitutes an effective amount. Further, several divided dosages, as well as staggered dosages, can be administered daily or sequentially, or the dose can be continuously infused, or can be a bolus injection. Further, the dosages of the compound(s) of the present disclosure can be proportionally increased or decreased as indicated by the exigencies of the therapeutic or prophylactic situation. An effective amount of the compounds of the present disclosure, sufficient for achieving a therapeutic effect, may range from about 0.000001 mg per kilogram body weight per day to about 10,000 mg per kilogram body weight per day. In some embodiments, the dosage ranges are from about 0.0001 mg per kilogram body weight per day to about 100 mg per kilogram body weight per day. The compounds disclosed herein can also be administered in combination with each other, or with one or more additional therapeutic compounds.
[0078] The term “sample”, as used herein, generally refers to a biological sample obtained or derived from a source of interest, as described herein. In some embodiments, a source of interest comprises an organism, such as an animal or human. In some embodiments, a biological sample comprises biological tissue or fluid. In some embodiments, a biological sample may be or comprise bone marrow; whole blood; plasma; serum; blood cells; platelets; ascites; tissue or fine needle biopsy samples; cell-containing body fluids; free floating nucleicacids; sputum; saliva; urine; cerebrospinal fluid, peritoneal fluid; pleural fluid; pancreatic cyst fluid; feces; lymph; gynecological fluids; skin swabs; vaginal swabs; oral swabs; nasal swabs; washings or lavages such as a ductal lavages or bronchoalveolar lavages; aspirates; scrapings; bone marrow specimens; tissue biopsy specimens; surgical specimens; amniotic fluid; semen; bile; gastric juice; breast milk; synovial fluid; pericardial fluid; otorrhea; ascitic fluid; pus; prostatic secretions; aqueous or vitreous humor; endolymph; perilymph; colonic lavages; intestinal secretions; hepatic secretions; intracranial fluid, other body fluids, secretions, and / or excretions; and / or cells therefrom, etc. In some embodiments, a biological sample is or comprises cells obtained from an individual. In some embodiments, a sample is a “primary sample” obtained directly from a source of interest by any appropriate method. For example, in some embodiments, a primary biological sample is obtained by methods selected from the group consisting of biopsy (e.g., fine needle aspiration, tissue biopsy, bone marrow biopsy, laparoscopic biopsy, liquid biopsy, punch biopsy, swabbing, lumbar puncture, etc.), surgery, collection of body fluid (e.g., blood, lymph, feces etc.), etc. In some embodiments, as will be clear from context, the term “sample” refers to a preparation that is obtained by processing (e.g., by removing one or more components of and / or by adding one or more agents to) a primary sample. For example, filtering using a semi-permeable membrane. Such a “processed sample” may comprise, for example nucleic acids or proteins extracted from a sample or obtained by subjecting a primary sample to techniques such as amplification or reverse transcription of RNA, isolation and / or purification of certain components, etc.
[0079] In some embodiments, the biological sample used for determining the level of one or more small non-coding RNA biomarkers is a sample containing circulating small ncRNAs, e.g., extracellular small ncRNAs. Extracellular small ncRNAs freely circulate in a wide range of biological materials, including bodily fluids, such as fluids from the circulatory system, e.g., a blood sample or a lymph sample, or from another bodily fluid such as urine or saliva. Accordingly, in some embodiments, the biological sample used for determining the level of one or more small ncRNA biomarkers is a bodily fluid, for example, blood, fractions thereof, serum, plasma, urine, saliva, tears, sweat, semen, vaginal secretions, lymph, bronchial secretions, CSF, whole blood, stool, interstitial fluid, synovial fluid, gastric acid, sebum, mucus, bile, etc. In some embodiments, the sample is a sample that is obtained non-invasively, such as a stool sample. In some embodiments, the sample is a serum sample from a human.
[0080] In some embodiments, any of the methods disclosed herein comprise using a small volume sample. In some embodiments, the methods disclosed comprise isolating total RNA or small RNA, e.g., small non-coding RNA, and / or amplifying total or small RNA in asample of no more than about 20 microliters of sample, 40 microliters of sample, 80 microliters of sample, 100 microliters of sample, 200 microliters of sample, 300 microliters of sample, 400 microliters of sample, 500 microliters of sample, 600 microliters of sample, 700 microliters of sample, 800 microliters of sample, 900 microliters of sample, 1 milliliter of sample, 1.1 milliliters of sample, 1.2 milliliters of samples, 1.3 milliliters of samples, 1.4 milliliters of samples, 1.5 milliliters of samples, 1.6 milliliters of samples, 1.7 milliliters of samples, 1.8 milliliters of samples, 1.9 milliliters of samples, 2.0 milliliters of sample. In some embodiments, the sample size is from about 25 microliters to about 2 milliliters of liquid sample in the form of subject plasma, whole blood or serum.
[0081] In some embodiments, the sample has a volume of about 0.001 mL to about 100 mL. In some embodiments, the sample has a volume of about 0.001 mL to about 0.01 mL, about 0.001 mL to about 0.1 mL, about 0.001 mL to about 0.5 mL, about 0.001 mL to about 1 mL, about 0.001 mL to about 10 mL, about 0.001 mL to about 100 mL, about 0.01 mL to about 0.1 mL, about 0.01 mL to about 0.5 mL, about 0.01 mL to about 1 mL, about 0.01 mL to about 10 mL, about 0.01 mL to about 100 mL, about 0.1 mL to about 0.5 mL, about 0.1 mL to about 1 mL, about 0.1 mL to about 10 mL, about 0.1 mL to about 100 mL, about 0.5 mL to about 1 mL, about 0.5 mL to about 10 mL, about 0.5 mL to about 100 mL, about 1 mL to about 10 mL, about 1 mL to about 100 mL, or about 10 mL to about 100 mL. In some embodiments, the sample has a volume of about 0.001 mL, about 0.01 mL, about 0.1 mL, about 0.5 mL, about 1 mL, about 10 mL, or about 100 mL. In some embodiments, the sample has a volume of at least about 0.001 mL, about 0.01 mL, about 0.1 mL, about 0.5 mL, about 1 mL, or about 10 mL. In some embodiments, the sample has a volume of at most about 0.01 mL, about 0.1 mL, about 0.5 mL, about 1 mL, about 10 mL, or about 100 mL.
[0082] Circulating small non-coding RNAs can include small non-coding RNAs in cells, extracellular small non-coding RNAs in microvesicles, in exosomes and extracellular small non-coding RNAs that are not associated with cells or microvesicles (extracellular, non- vesicular small non-coding RNA). In some embodiments, the biological sample used for determining the level of one or more small non-coding RNA biomarkers (e.g., a sample containing circulating small non-coding RNA) may contain cells. In other embodiments, the biological sample may be free or substantially free of cells (e.g., a plasma or serum sample). In some embodiments, a sample containing circulating small non-coding RNAs, e.g., extracellular small non-coding RNAs, is a blood-derived sample. Blood-derived sample types may include, e.g., a plasma sample, a serum sample, a blood sample, etc. In other embodiments, a sample containing circulating small non-coding RNAs is a lymph sample. Circulating small non-codingRNAs are also found in urine and saliva, and biological samples derived from these sources are likewise suitable for determining the level of one or more small non-coding RNA biomarkers.
[0083] In some embodiments, biological samples may be subjected to one or more processing operations as part of methods and systems as described herein. A processing operation may be carried out to isolate a substance from the biological sample, to purify the biological sample, to separate the biological sample into one or more fractions for further use or processing, to quantitate the amount of a substance, to detect the presence or absence or one or more substances, to transform or modify a substance for further downstream processing or analysis, or any combination thereof. For example, enzymes may be added to digest protein and remove contamination, or to inactivate nucleases that might otherwise degrade nucleic acids (RNA or DNA) during purification. The one or more processing operations may immediately follow sample collection, may be immediately prior to another processing or assaying operation, or may be carried out contemporaneously with another processing or assaying operation. The processing operation(s) may also be carried out on a sample or an intermediate in the process that has been appropriately stored for a designated amount of time. Any number of suitable processing operations may be performed on the biological sample or part thereof.
[0084] Processing operations as described herein may include, but are not limited to, immunoassays, enzyme-linked immunosorbent assays (ELISA), radioimmunoassays (RIA), ligand binding assays, functional assays, enzymatic assays, enzymatic treatments (e.g., with kinases, phosphatases, ligases, transcriptases, reverse transcriptases), enzymatic digestions (e.g., nucleases), spectroscopic assays (e.g., UV-vis spectroscopy, Fourier transform infrared spectroscopy, circular dichroism spectroscopy) spectrophotometric assays (e.g., ultraviolet- visible light spectrophotometry), immunoprecipitations (IP), sequencing reactions, electrophoresis, chromatography, enrichments, pull-downs, and mass spectrometry (MS). In some embodiments, a method as described herein comprises not performing one or more processing operations.
[0085] In some embodiments, the biological sample is subjected to one or more immunoprecipitation (IP) reactions, such as an in vitro or in vivo crosslinking and IP reaction (CLIP). The one or more IP reactions may enrich or pull-down one or more substances of interest. In some embodiments, an IP processing operation includes a cross-linking operation to covalently link two or more different substances (e.g., to link protein and DNA or protein and RNA). Immunoprecipitation of the cross-linked substances may provide an indication of biological substances that are associated with one another and / or may be used to enrich forspecific substances that are known or suspected to interact with one another. In an example, one or more RNAs of interest are cross-linked to one or more corresponding proteins. IP of the proteins cross-linked with the RNAs allows for subsequent isolation and downstream processing of the RNAs of interest. In another example, an IP reaction may be carried out using antibodies specific for an RNA modification of interest, e.g., an adenosine modification (such as m6A, mxA, alternative polyadenylation, or adenosine-to-inosine RNA editing), a uridine modification (such as conversion to pseudouridine), or other RNA modifications as alluded to above.
[0086] In some embodiments, the sample is subjected to one or more isolation operations. Isolation operations may target a general class of molecules (e.g., nucleic acids, such as RNAs) or a specific molecule (e.g., a specific annotated RNA molecule).
[0087] In some embodiments, a processing operation may comprise adding (e.g., spiking in) one or more substances. The one or more spike-in substances may be for any suitable purpose, including but not limited to, quality control, enrichment of target species, the act of depleting non-target species, or any combination thereof. In some embodiments, a spike-in substance may comprise a synthetic biomolecule (e.g., nucleic acid, such as an RNA or a modified RNA, i.e., an RNA containing base modifications as alluded to above) corresponding to a target biomolecule. In some embodiments, the spike-in substance may comprise an endogenous or exogenous biomolecule. The spike-in molecule may be selected based on any appropriate property such as abundance or relative abundance, origin, sequence or part thereof, global or local structure (such as secondary or tertiary structure), or any combination thereof. Quantification of a spike-in substance following one or more downstream processing operations may be used for quality control as well as quantity control.
[0088] In some embodiments, a sample is processed or analyzed by a third party. For example, one party may conduct a purification or enrichment processing operation on a sample. The purified or enriched sample may then be subjected to a subsequent processing operation (e.g., a quantitation) by a different party.
[0089] In some embodiments, biological samples containing or suspected of containing one or more RNAs are subjected to one or more processing operations to remove or change RNA modifications that may inhibit further downstream processing (e.g., subsequent reverse transcription). Alternatively or additionally, chemical modifications may be removed or retained to determine the presence or absence of a relationship between chemical modifications and a disease state or any pathological state, e.g., a cancer state of a subject or population of subjects. Chemically modified RNA bases may include, without limitation, N6-methyladenosine(m6A), inosine (I), 5 -methylcytosine (m5C), pseudouridine ( ), 5-hydroxymethylcytosine (hm5C), bri-methyladenosine (nriA), or 7-methylguanosine (m7G). For example, biological samples which contain or are suspected of contained methylated RNAs (e.g, comprising m6A, m5C, hm5C, nriA, or m7G) may be treated with one or more demethylation enzymes (e.g, AlkB) to remove alkyl groups which may interfere with downstream processing (e.g., by reverse transcriptases).
[0090] In some embodiments, biological samples containing or suspected of containing RNAs are subjected to one or more processing operations to facilitate downstream processing operations (e.g., isolation). In some embodiments, a sample containing or suspected of containing RNAs may be treated with one or more enzymes, cofactors, and / or other reagents to bring about an end-repair process. The sample containing or suspected of containing one or more RNAs is subjected to treatment with a polynucleotide kinase (PNK) to ensure that end- modified RNA species having a 3'-phosphate group are not lost to further analysis. That is, PNK enzymes dephosphorylate 3'-phosphate RNA species and thus allow for subsequent polyadenylation and inclusion in downstream processing steps.
[0091] The terms “treating” or “treatment” or “ treat”, as used herein, generally refer to both 1) therapeutic measures that cure, slow down, lessen symptoms of, and / or halt progression of a diagnosed pathologic condition or disorder and 2) prophylactic or preventative measures that prevent or slow the development of a targeted pathologic condition or disorder. Thus those in need of treatment include those already diagnosed with the disorder; those prone to have the disorder; and those in whom the disorder is to be prevented. In some embodiments, a subject is successfully “treated” according to the methods of the present disclosure if the patient shows one or more of the following: a reduction in the number of and / or complete absence of cancer cells; a reduction in the tumor size; an inhibition of tumor growth; inhibition of and / or an absence of cancer cell infiltration into peripheral organs including the spread of cancer cells into soft tissue and bone; inhibition of and / or an absence of tumor or cancer cell metastasis; inhibition and / or an absence of cancer growth; relief of one or more symptoms associated with the specific cancer; reduced morbidity and mortality; improvement in quality of life; reduction in tumorigenicity; reduction in the number or frequency of cancer stem cells; or some combination of such effects.
[0092] The term "tumor", as used herein, generally refers to all neoplastic cell growth and proliferation, whether malignant or benign, and all pre-cancerous and cancerouscells and tissues. The term also comprises a cancerous cell or a cancerous tissue that is residual after a cancer treatment.
[0093] The term "tumor sample", as used herein, generally refers to a sample comprising tumor material obtained from a cancer patient. The term encompasses tumor tissue samples, for example, tissue obtained by surgical resection and tissue obtained by biopsy, such as for example, a core biopsy or a fine needle biopsy. In a particular embodiment, the tumor sample is a fixed, wax-embedded tissue sample, such as a formalin-fixed, paraffin-embedded tissue sample. Additionally, the term "tumor sample" encompasses a sample comprising tumor cells obtained from sites other than the primary tumor, e.g., circulating tumor cells. The term also encompasses cells that are the progeny of the patient's tumor cells, e.g. cell culture samples derived from primary tumor cells or circulating tumor cells. The term further encompasses samples that may comprise protein or nucleic acid material shed from tumor cells in vivo, e.g., bone marrow, blood, plasma, serum, and the like. The term also encompasses samples that have been enriched for tumor cells or otherwise manipulated after their procurement and samples comprising polynucleotides and / or polypeptides that are obtained from a patient' s tumor material. The term also encompasses samples from a patient that has received a cancer treatment.
[0094] RNA sequencing can be used to identify a category of cancer-associated, small non-coding RNAs, termed orphan non-coding RNAs (oncRNAs). OncRNAs can be actively secreted from living cancer cells and are stable and abundant in the blood of cancer patients. A catalog of hundreds of thousands of oncRNAs and thousands of patient oncRNA profiles, spanning all major cancer types, can be used to diagnose cancer or minimal residual disease (MRD). When combined with proprietary artificial intelligence technology, the systems and methods described herein can have increased sensitivity and specificity, as well as the ability to reveal dynamic changes in the biology of a patient’s tumor over time. The systems and methods described herein can detect or predict a presence or an absence of MRD in a patient. The systems and methods described herein can be used across multiple cancer care settings such as screening and early detection, monitoring, molecular residual disease and therapy selection.
[0095] The systems and methods described herein can be used to detect cancer at the earliest stages and the smallest tumor sizes, as well as minimal residual disease (MRD), using a standard blood sample. In some cases, a tumor size can be predicted from an oncRNA signature with a correlation coefficient to an actual tumor size of about 0.45 to about 0.55. In some cases, a metastasis status can be predicted with a sensitivity of 80% and specificity of 81%. In somecases, stage I cancer can be distinguished with a sensitivity about 57% at 90% specificity. In some cases, stage II cancer can be distinguished with a sensitivity about 64% at 90% specificity. In some cases, stage III cancer can be distinguished with a sensitivity about 74% at 90% specificity. In some cases, stage II cancer can be distinguished with a sensitivity about 83% at 90% specificity. In some cases, a T score of T1 can be distinguished with a sensitivity about 54% at 90% specificity. In some cases, a T score of T2 can be distinguished with a sensitivity about 69% at 90% specificity. In some cases, a T score of T3 can be distinguished with a sensitivity about 75% at 90% specificity. In some cases, a T score of T1 can be distinguished with a sensitivity about 80% at 90% specificity. Overall sensitivity across all cancer stages can be about 64% sensitivity at 90% specificity, far exceeding DNA-based liquid biopsy performance which falls short in cancer detection due to low shedding rates and technical barriers.
[0096] The systems and methods described herein can identify de novo expression profiles from tumor samples, enabling the discovery and use of previously unknown or unannotated RNAs (including non-coding RNAs).Machine Learning Model
[0097] The present disclosure provides a trained machine learning algorithm that uses variational inference to model an underlying distribution of small RNAs, as described in Karimzadeh, M., Momen-Roknabadi, A., Cavazos, T.B. et al. Deep generative Al models analyzing circulating orphan non-coding RNAs enable detection of early-stage lung cancer. Nat Commun 15, 10090 (2024). https: / / doi.org / 10.1038 / §41 67-024-53851-9. In some cases, the underlying distribution of small RNAs comprise oncRNAs.
[0098] In some cases, the trained machine learning algorithm comprises an autoencoder. In some cases, the autoencoder is a variational autoencoder. In some cases, the machine learning algorithm comprises a neural network. In some cases, the neural network is a regressor neural network. In some cases, the neural network is a multi-head neural network. In some cases, the neural network comprises a regressor neural network and a multi-head neural network.
[0099] In some cases, the trained machine learning algorithm can predict a presence or an absence of residual cancer cells in a subject after said subject has received a cancer therapy. In some cases, the cancer therapy is a curative intent cancer therapy. In some cases, the cancer therapy is a neoadjuvant therapy. In some cases, the cancer therapy is an adjuvanttherapy. In some cases, predicting the presence or the absence of the residual cancer cells in a subject after said subject has received a cancer therapy can be used to evaluate the efficacy of the cancer therapy. In some cases, predicting the absence of the residual cancer cells can be indicative of an absence of minimal residual disease (MRD). In some cases, predicting the presence or the absence of the residual cancer cells can be used to determine an appropriate cancer therapy.
[0100] In some cases, the trained machine learning algorithm can predict TNM staging scores, comprising a tumor score (T score), a node score (N score), and a metastasis score (M score). In some cases, a T score can describe the size and extent of a tumor, and can be categorized by a score of T0-T4, with a higher score being indicative of a larger or more advanced tumor. In some cases, a score of TO is indicative of a pre-cancerous condition, such as ductal carcinoma in situ (DCIS). In some cases, the N score can describe whether a cancer has spread to a nearby lymph node, and can be categorized by a score of N0-N3, with a higher score being indicative of a cancer spreading to one or more lymph nodes. In some cases, the N score can be used to predict a presence or an absence of infiltration of a cancer into a lymph node (e.g., lymphovasular invasion). In some cases, the M score can describe whether a cancer has spread from a primary tumor site to a different tumor site (e.g., a tumor has metastasized), and can be categorized as MO (no metastasis) or Ml (metastasis present). In some cases, the TNM staging score can be used to determine a progression of a cancer, a stage of a cancer, a subtype of a cancer, or any combination thereof in a subject. In some cases, the TNM staging score can be used to determine an appropriate cancer therapy for said subject.
[0101] In some cases, the trained machine learning algorithm can predict a tumor burden or a tumor size in a subject. In some cases, the predicted tumor burden or tumor size can be used to determine an appropriate cancer therapy for said subject. In some cases, the predicted tumor burden or tumor size can be used to stratify a patient for tumor severity. In some cases, the stratification can be used to determine an appropriate cancer therapy for said patient.
[0102] In some cases, the trained machine learning algorithm can predict a site of tumor origin. In some cases, the predicted site of tumor origin can be used to determine an appropriate cancer therapy.
[0103] As illustrated in FIG. 3A, the systems and methods described herein can be used to indicate cancer presence, residual cancer cell presence, or tumor absence in a liquid biopsy sample obtained from a patient (i.e., serum, plasma, etc.) 306. The patient’s individual oncRNA profile (small serum RNA profile) 305 can be determined from the patient sample. Thepatient’s oncRNA profile 305 can be input into a generative Al model 303 that compares the patient’s oncRNA profile 305 with a breast cancer oncRNA fingerprint 302 to make one or more oncRNA-based predictions 304.
[0104] The breast cancer oncRNA fingerprint 302 may be obtained through an oncRNA discovery model 301. In some embodiments, an independent non-cancer serum reference cohort can be received as input by the oncRNA discovery model 301. In some embodiments, samples from patients with or without residual cancer cells are input into the oncRNA discovery model 301.
[0105] In some embodiments, the oncRNA discovery model 301 can generate a cancer oncRNA fingerprint 302. In some embodiments, the cancer oncRNA fingerprint 302 can be received as input by a generative Al model 100. In some embodiments, patient serum or plasma 306 can be processed to generate a small RNA profile 305. In some embodiments, the generative Al model 100 can additionally receive input data comprising the patient oncRNA profile 305. In some embodiments, the generative Al model 100 can process the input cancer oncRNA fingerprint data 302 and the patient oncRNA profile 305 to produce one or more oncRNA-based predictions 304.
[0106] FIG. 1 shows a detailed view of the generative Al model 100. A generative Al model similar to that disclosed in “Deep generative Al models analyzing circulating orphan non-coding RNAs enable accurate detection of early-stage non-small cell lung cancer. Mehran Karimzadeh et al. medRxiv 2024.04.09.24304531; doi: https. / / doi,org / 10.1101 / 2024,04,09,2430453 i’’ can be used. This paper is herein incorporated by reference in entirety. The model 100 can distinguish three groups of samples during training: 1) samples from individuals without any diagnosis of cancer or residual cancer (healthy controls), 2) samples with individuals with a diagnosis of cancer or residual cancer, and 3) samples from individuals with a diagnosis of cancer. The model 100 can estimate the probability of a patient sample belonging to either of these three groups. In some cases, the model provides three probability scores (one for each of the groups described above).
[0107] FIG. 1 illustrates an example of a method of using the trained machine learning algorithm (or generative Al model) to perform a multi -class prediction of a metastatic status, a nodal status, a tumor grade, a tumor burden status, or any combination thereof, based at least in part on an oncRNA count from a patient sample. FIG. 1 illustrates inputting a count matrix of oncRNAs 102 and a count matrix of endogenous RNAs 104 into the trained machine learning algorithm 100. The count matrix of the oncRNAs 102 are processed by an oncRNAencoder 106 to determine a variational inference of the oncRNAs 108. The count matrix of the endogenous RNAs 104 are processed by a library encoder 110 to determine a variational inference of the endogenous RNAs 112. In some cases, the oncRNA encoder 104 and the library encoder 110 are variational autoencoders. FIG. 1 illustrates performing generative sampling 114 by the trained machine learning algorithm 100 on the variational inference of the oncRNAs 108 and the variational inference of the endogenous RNAs 112. The generative sampling 114 can be input to a decoder 116 to generate a reconstruction loss 118. The generative sampling 114 can be input to a regressor network 120 to generate a tumor burden prediction 122, based at least in part on the oncRNA count matrix 102 and the endogenous RNA count matrix 104. The generative sampling 114 can be input to a multi -head neural network 124 to output a prediction of a metastatic status 126, a nodal status 128, and a tumor grade 130.
[0108] In another example, FIG. 2 illustrates a method of training a machine learning algorithm 200 to predict a tumor burden from clinical characteristics of a patient tumor sample. The clinical characteristics comprising a cancer stage 202, a T stage 204, a node invasion score 206, a metastasis score 208, a tumor grade 210, a tumor size 212, a cancer site 214, or any combination thereof, can be input to a machine learning algorithm 200. The machine learning model is trained to generate a continuous measurement of tumor burden for the patient sample. The trained machine learning model can output a computational low-dimensional continuous representation of multiple clinical variables from the clinical characteristics of the tumor sample (e.g., a tumor burden measurement) 122. This tumor burden measurement 122 has sufficient information to provide a predicted cancer stage 216, a predicted T stage 218, a predicted node invasion score 128, a predicted metastasis score 126, a predicted tumor grade 130, a predicted tumor size 226, a predicted cancer site 228, or any combination thereof. The tumor burden measurement 122 can be used as a target label for oncRNA-based machine learning models, as described herein. In some cases, a target label is an output of the trained machine learning model.
[0109] In some cases, the trained learning algorithm can output a tumor burden measurement 122, based at least in part on an input of a count matrix of oncRNAs 102. In some cases, the trained machine learning algorithm can output clinical characteristics comprising a predicted cancer stage 216, a predicted T stage 218, a predicted node invasion score 128, a predicted metastasis score 126, a predicted tumor grade 130, a predicted tumor size 226, a predicted cancer site 228, or any combination thereof, based at least in part on an input of a count matrix of oncRNAs 102 and / or the tumor burden measurement 122.
[0110] In some cases, the systems and methods described herein can predict cancer v. non-cancer with a false discovery rate less than or equal to 0.5. In some cases, the systems and methods described herein can predict cancer v. non-cancer with a false discovery rate less than or equal to 0.2. In some cases, the systems and methods described herein can predict cancer v. non-cancer with a false discovery rate less than or equal to 0.1. In some cases, the systems and methods described herein can predict cancer v. non-cancer with a false discovery rate less than or equal to 0.05.EXAMPLESTraining the Machine Learning Algorithm
[0111] The present disclosure provides a trained machine learning algorithm to perform a multi -class prediction of a metastatic status, a nodal status, a tumor grade, a tumor burden status, or any combination thereof, based at least in part on an oncRNA count from a patient sample, as described herein.
[0112] The trained machine learning algorithm was trained based on principal component analysis (PCA), non-negative matrix factorization (NMF), and a customized neural architecture leveraging variational Bayes for learning computational measures of tumor burden. These burden measures were used to train a machine learning model on cell-free oncRNA from a dataset for serum of breast cancer patients. A data augmentation strategy was used to improve model robustness with respect to RNA yield and sequencing depth by taking samples from the sequencing reads without replacement to generate in silico versions of the data at lower sequencing depths. During model training, these in silico generated augmented datasets were trained on by ensuring no leakage of different augmentations of each sample into training and validation splits of the cross validation.
[0113] The trained machine learning algorithm was generated by training a machine learning algorithm on a dataset of 719 treatment-naive patient serum samples. The treatment- naive patient serum samples were derived from a set of patients across multiple demographics, cancer types, and cancer stages. The data set comprise TNM staging scores, comprising a tumor score (T score), a node score (N score), and a metastasis score (M score). The T score can describe the size and extent of a tumor, and can be categorized by a score of T1-T4, with a higher score being indicative of a larger or more advanced tumor. The N score can describe whether a cancer has spread to a nearby lymph node, and can be categorized by a score of NONS, with a higher score being indicative of a cancer spreading to one or more lymph nodes. The M score can describe whether a cancer has spread from a primary tumor site to a different tumorsite (e.g., a tumor has metastasized), and can be categorized as MO (no metastasis) or Ml (metastasis present). The machine learning model was trained to use a cell-free oncRNA expression level from a patient sample to predict the clinical characteristics, as described in FIG. 2, of the patient’s tumor. Validation of the trained machine learning algorithm was performed through a 10-fold cross-validation with five repeats (bags), providing a comprehensive assessment of model stability and reliability. Table 1 shows a count of patients by T score in the training dataset.Table 1: Frequency of patient T scores in the training dataset.Performance of the Trained Machine Learning Algorithm
[0114] The trained learning algorithm generates prediction of a metastatic status, a nodal status, a tumor grade, a tumor burden status, or any combination thereof, based at least in part on an oncRNA count from a patient sample, using inputs comprising tumor size (such as a largest diameter of a radiographic image), functional tumor volume, and TNM staging scores
[0115] FIG. 3B shows the performance of the trained machine learning algorithm to predict tumor size. The trained machine learning algorithm output a predicted tumor size from patients diagnosed with breast and lung cancer across cancer stages. The predicted tumor size was transformed by Rank-Based Inverse Normal Transformation, to generate a transformed tumor size prediction. The transformed tumor size prediction was correlated to an actual tumor size of the patients diagnosed with breast and lung cancer, using Pearson’s correlation, and resulted in a correlation coefficient of R = 0.52, with a p values of less than 2.2e-16, and a coefficient of determination of R2= 0.27. These results indicate that the machine learning algorithm predicted tumor size, with a significantly positive correlation.
[0116] FIGs. 4A-4C show the performance of the trained machine learning algorithm to predict tumor burden and tumor grade. Predicted tumor size from the machine learning algorithm was compared to actual T scores from patients diagnosed with breast and lung cancer. FIG. 4A shows predicted tumor size (y-axis) plotted against actual tumor grade (Tscore) (x-axis). The results show a general increase in predicted size of the tumor with higher tumor grade. FIG. 4B shows actual tumor size (y-axis) plotted against actual tumor grade (T score) (x-axis). The results show a comparable trend of increase in tumor size with higher tumor grade. FIG. 4C shows a correlation between predicted tumor burden and tumor stage, with predicted tumor size (y-axis) plotted against actual tumor stage (x-axis). The results provide a systematic increase in tumor size prediction with increases in actual tumor stage, suggesting that the trained machine learning model can accurately predict tumor size or tumor burden.
[0117] FIG. 5 shows a concordance between a metastatic status predicted by the trained machine learning algorithm and actual metastatic status. Metastasis status was obtained by pathology reports across patients diagnosed with breast or lung cancer. FIG. 4 shows that trained machine learning algorithm predicted M status with a sensitivity of 80% and specificity of 81%.
[0118] FIG. 6 provides a concordance index (y-axis) as a measure of the performance of a multi-variate Cox proportional hazards model in predicting patient overall survival from a tumor burden output (the first three bars from left to right) of the trained machine learning model, compared to actual tumor size (the fourth bar from left to right), in all cancer types (n=707), breast cancer (n=196), colorectal cancer (n=l 14), and lung cancer (n=397). These data show that the machine learning algorithm had high concordance in predicting overall survival from a tumor burden prediction, as compared to actual tumor size.
[0119] The trained machine learning model was tested on a held-out dataset of a cancer patient oncRNA dataset and a control oncRNA dataset. FIG. 7A shows a significant, positive correlation of predicted tumor burden (x-axis) with actual tumor size (y-axis). FIG. 7B shows a significant, positive correlation of predicted tumor burden (x-axis) with computational tumor burden (y-axis). FIG. 7C shows the area under the receiver-operating characteristic curve for distinguishing cancer oncRNA samples from control oncRNA samples, with a sensitivity of 64.1% at 90% specificity using the trained machine learning algorithm. FIG. 7D shows the sensitivity for distinguishing tumors of different stages at 90% specificity using the trained machine learning algorithm. FIG. 7E shows the sensitivity for distinguishing tumors of different T score at 90% specificity using the trained machine learning algorithm. For FIGs. 7D-7E the numbers in parentheses on the x-axis indicate the number of samples at each stage.Prognosis of TNBC using a oncRNA signature
[0120] A study was performed to investigate the capability of an oncRNA signature to predict a risk of recurrence of residual disease (RD) in triple-negative breast cancer (TNBC) patients with RD. The study population included patients with stage I-III TNBC (less than 10% expression of estrogen receptor (ER) and progesterone receptor (PR), and HER-2 negative) with RD. The study had access to end-of-treatment (EOT) serum samples for patients who were enrolled in an IRB-approved multisite prospective cohort study (P.R.O.G.E.C.T., NCT02302742). EOT samples were collected 14-180 days after completion of all curative treatment (local / systemic). Small RNAs isolated from EOT serum were sequenced at average depth of 76.5 ±12.5 million 50 bp single-end reads and annotated using a bioinformatics pipeline to identify oncRNAs (Table 2). Cancer risk scores were generated using an oncRNA-based tumor detection model trained on 451 treatment-naive breast cancer samples and 470 samples from individuals without known cancer diagnosis.
[0121] The study cohort was divided into training and testing sets of equal size. Score cutpoint for high vs low recurrence risk was determined in the testing set through ROC analysis. The training and testing sets were balanced for clinico-pathologic characteristics, treatment, and pathologic response. The training set (n=39) was used to define a oncRNA recurrence risk score cutpoint, which was applied to the testing set (n=40). The impact of an EOT oncRNA risk category on event-free survival (EFS) and overall survival (OS) was estimated by Kaplan-Meier method and compared between groups by log-rank test, followed by Cox regression modeling. Residual cancer burden (RCB) was determined according to the classification by Symmans et al. (Symmans, W. F., Hatzis, C., Sotiriou, C., Andre, F., Peintinger, F., Regitnig, P., ... & Pusztai, L. (2010). Genomic index of sensitivity to endocrine therapy for breast cancer. Journal of clinical oncology, 28(27), 4111-4119). FIG. 8 provides the demographics, BRCA1 / 2 mutation status, T stage, nodal status, TNM stage, RCB class, adjuvant radiation status, and adjuvant chemotherapy status of the study cohort.
[0122] The results of the study revealed that oncRNA isolation and score generation was successful for 79 out of 80 TNBC patients with RD with available EOT serum samples. RCB class distribution was: RCB 1=27%, RCB 11=49%, RCB 111=18%. The median age of the cohort was 48 years and 39% had node-positive disease. Fifteen of forty patients (38%) in the testing set were classified as oncRNA high-risk. oncRNA high-risk status was associated with higher T stage (p=0.018) (FIG. 8). Rates of oncRNA high-risk status in RCB I, II, and III groups were 25%, 41%, and 57%, respectively.
[0123] In the testing set, at a median follow-up of 29 months, oncRNA high-risk status was associated with lower event-free survival (EFS) (FIG. 9 and FIG. 10A) and overall survival (OS) (FIG. 9 and FIG. 10B).
[0124] In multivariable analysis including oncRNA risk status, T stage, nodal status, and RCB class, oncRNA high risk status and RCB class retained significant association with EFS (high-risk oncRNA status (HR) 7.70, 95% CI 1.33-44.64, p=0.023 and HR 40.33, 95% CI 2.67-608.99, p=0.008, respectively) and OS (HR 7.99, 95% CI 1.36-47.00, p=0.022 and HR 28.59, 95% CI 2.47-601.93, p=0.009, respectively).
[0125] Among patients with RCB classes I-II, oncRNA low-risk status was associated with better outcomes: 3 year EFS 94% vs 66% (HR 6.31, 95% CI 0.65-60.88, p=0.068) (FIG. 10C) and 3y OS 93% vs 76% (HR 7.30, 95% CI 0.75-70.81, p=0.045) for low vs high risk (FIGs. 10B). Patients with RCB III had suboptimal outcomes regardless of oncRNA levels (7 / 7 of patients with RCB III in the testing cohort had an EFS event).
[0126] These results provide that oncRNA liquid biopsy recurrence risk assay had good technical success in EOT samples from TNBC patients with RD and was highly prognostic, with almost half of patients in the oncRNA high-risk group suffering an EFS event by 3 years. oncRNA risk score has the potential to provide prognostic utility complementary to clinicopathologic characteristics in patients with TNBC. In patients with RCB I-II, oncRNA risk model provided further prognostic utility, as patients with oncRNA low-risk status had excellent outcomes (3y EFS / OS >93%). These patients could potentially be spared further adjuvant therapy intensification.Predicting Long-Term Survival in Neoadjuvant Breast Cancer Therapy using oncRNAs
[0127] Pathologic complete response (pCR) is a valuable metric for predicting survival following neoadjuvant therapy (NAT) in breast cancer. Blood-based biomarkers may further refine the prediction of treatment outcomes to inform follow-up treatment decisions. Orphan non-coding RNAs (oncRNAs) are a category of small RNAs (smRNAs) that are frequently detected in cancer and largely absent in non-cancerous tissues. Tumor-naive oncRNA-based liquid biopsy assays were used in a study for predicting distant recurrence-free survival (DRFS) in breast cancer patients using a standard blood draw.
[0128] The study included women diagnosed with stage II / III breast cancer with MammaPrint “High Risk” of recurrence. The cohort included 538 patients with available serum following NAT and prior to surgery (timepoint T3). Small RNA was isolated from 0.8 mL of patient serum, and sequences libraries were generated from the isolated small RNAs. The libraries were sequenced at an average depth of 50 million 100-bp single end reads. A trained machine learning algorithm, as described herein, was developed for predicting an oncRNA riskscore using a catalog of tumor-derived oncRNAs identified in The Cancer Genome Atlas (Table 2). The machine learning algorithm was trained on oncRNA profiles from an independent cohort of 719 serum samples collected at diagnosis from individuals with cancer, using tumor size as the outcome of interest, as described herein. oncRNAs were used to assess a risk of cancer recurrence post-NAT (T3). The risk of cancer recurrence was associated with distance relapse- free survival (DRFS) using cutpoints established for two pCR groups comprising a high-risk group (n=40) and a low-risk group (n=498). Cutpoints were verified through cross-validation and almost all folds converged on a single value for each pCR group.
[0129] FIG. 11 illustrates the study design. The study used adaptive randomization to assess efficacy of drugs in sequence with standard chemotherapy, to identify treatments for patients based on molecular characteristics and to achieve a high patient pCR rate. FIG. 12 shows the study demographics. Hormone receptor (HR) status was missing for 3 patients, residual cancer burned (RCB) was missing for 2 patients, and T stage was missing for 9 patients. FIG. 13 shows a breakdown of patient receptor subtypes by treatment arm. The study cohort consisted of 1 control arm and 9 experimental treatment arms. The experimental treatment arms comprised of 3 completed arms (AMG386, Ganitumab, and Ganetespib) and 6 graduated arms (Neratinib, VC, TDM1 / P, Pertuzumab, MK2206, and Pembro). Patients were randomized based on ER and HER2 status, and MammaPrint score.
[0130] FIG. 14 provides data that indicate continuous oncRNA risk scores are predictive of distant recurrence in univariate and multivariate models. Hazards ratios (HR) and 95% Confidence Intervals (CI) from Cox models were reported with p-values from the Wald test. # indicates reported at timepoint T3 (post-NAT). f indicates continuous variables are normalized so that the unit of measure for the HR is per standard deviation. J indicates all significant variables from univariate analysis remain significant when combined in multivariate analysis.
[0131] FIGs. 15A-15D show Kaplan-Meier survival curves for predicting DRFS when stratifying by pCR alone (FIG. 15A), oncRNA risk alone (FIG. 15B), or pCR further stratified by oncRNA risk score group within patients that achieved pCR (FIG. 15C), and patients that did not achieve pCR (FIG. 15D). Through a Cox multivariate analysis, oncRNA predicted risk (high vs. low; HR = 4.1 [95% CI: 2.4-6.8]) was identified as an independent prognostic factor when combined with pCR. The continuous oncRNA predicted risk measure was also significant in a multivariate analysis with pCR (p=2.38xl0‘3).Computer Systems
[0132] The present disclosure provides computer systems that are programmed to implement methods of the disclosure. FIG. 16 shows a computer system 701 that isprogrammed or otherwise configured to, for example, implement a trained machine learning classifier to determine the cancer status a subject from small RNA sequencing read derived from a sample from the subject, plan a course of treatment for a subject based on small RNA sequencing reads derived from a sample from the subject, monitor the cancer status of a subject over time (e.g., before and after treatment) based on small RNA sequencing reads derived from a sample from the subject, measure the efficacy or tolerability of a therapeutic intervention for an individual based on small RNA sequencing reads derived from a sample from the subject, or process a biological sample as described herein.
[0133] The computer system 701 can regulate various aspects of analysis, calculation, and generation of the present disclosure, such as, for example, implement a trained machine learning classifier to determine the cancer status a subject on small RNA sequencing reads derived from a sample from the subject, planning a course of treatment for a subject based on small RNA sequencing reads derived from a sample from the subject, monitoring the cancer status of a subject over time (e.g., before and after treatment) based on small RNA sequencing reads derived from a sample from the subject, measuring the efficacy or tolerability of a therapeutic intervention for an individual based on small RNA sequencing reads derived from a sample from the subject, or processing a biological sample as described herein. The computer system 701 can be an electronic device of a user or a computer system that is remotely located with respect to the electronic device. The electronic device can be a mobile electronic device.
[0134] The computer system 701 includes a central processing unit (CPU, also “processor” and “computer processor” herein) 705, which can be a single core or multi core processor, or a plurality of processors for parallel processing. The computer system 701 also includes memory or memory location 110 (e.g., random-access memory, read-only memory, flash memory), electronic storage unit 715 (e.g., hard disk), communication interface 720 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 725, such as cache, other memory, data storage and / or electronic display adapters. The memory 110, storage unit 715, interface 720 and peripheral devices 725 are in communication with the CPU 705 through a communication bus (solid lines), such as a motherboard. The storage unit 715 can be a data storage unit (or data repository) for storing data. The computer system 701 can be operatively coupled to a computer network (“network”) 730 with the aid of the communication interface 720. The network 730 can be the Internet, an internet and / or extranet, or an intranet and / or extranet that is in communication with the Internet.
[0135] The network 710 In some embodiments is a telecommunication and / or data network. The network 730 can include one or more computer servers, which can enable distributed computing, such as cloud computing. For example, one or more computer servers may enable cloud computing over the network 730 (“the cloud”) to perform various aspects of analysis, calculation, and generation of the present disclosure, such as, for example, implementing a trained machine learning classifier to determine the cancer status a subject on small RNA sequencing reads derived from a sample from the subject, planning a course of treatment for a subject based on small RNA sequencing reads derived from a sample from the subject, monitoring the cancer status of a subject over time (e.g., before and after treatment) based on small RNA sequencing reads derived from a sample from the subject, measuring the efficacy or tolerability of a therapeutic intervention for an individual based on small RNA sequencing reads derived from a sample from the subject, or processing a biological sample as described herein. Such cloud computing may be provided by cloud computing platforms such as, for example, Amazon Web Services (AWS), Microsoft Azure, Google Cloud Platform, and IBM cloud. The network 730, In some embodiments with the aid of the computer system 701, can implement a peer-to-peer network, which may enable devices coupled to the computer system 701 to behave as a client or a server.
[0136] The CPU 705 may comprise one or more computer processors and / or one or more graphics processing units (GPUs). The CPU 705 can execute a sequence of machine- readable instructions, which can be embodied in a program or software. The instructions may be stored in a memory location, such as the memory 110. The instructions can be directed to the CPU 705, which can subsequently program or otherwise configure the CPU 705 to implement methods of the present disclosure. Examples of operations performed by the CPU 705 can include fetch, decode, execute, and writeback.
[0137] The CPU 705 can be part of a circuit, such as an integrated circuit. One or more other components of the system 701 can be included in the circuit. In some embodiments, the circuit is an application specific integrated circuit (ASIC).
[0138] The storage unit 715 can store files, such as drivers, libraries and saved programs. The storage unit 715 can store user data, e.g., user preferences and user programs. The computer system 701 in some embodiments can include one or more additional data storage units that are external to the computer system 701, such as located on a remote server that is in communication with the computer system 701 through an intranet or the Internet.
[0139] The computer system 701 can communicate with one or more remote computer systems through the network 730. For instance, the computer system 701 can communicate with a remote computer system of a user. Examples of remote computer systems include personal computers (e.g., portable PC), slate or tablet PC’s (e.g., Apple® iPad, Samsung® Galaxy Tab), telephones, Smart phones (e.g., Apple® iPhone, Android-enabled device, Blackberry®), or personal digital assistants. The user can access the computer system 701 via the network 730.
[0140] Methods as described herein can be implemented by way of machine (e.g., computer processor) executable code stored on an electronic storage location of the computer system 701, such as, for example, on the memory 110 or electronic storage unit 715. The machine executable or machine readable code can be provided in the form of software. During use, the code can be executed by the processor 705. In some embodiments, the code can be retrieved from the storage unit 715 and stored on the memory 110 for ready access by the processor 705. In some situations, the electronic storage unit 715 can be precluded, and machine-executable instructions are stored on memory 110.
[0141] The code can be pre-compiled and configured for use with a machine having a processer adapted to execute the code or can be compiled during runtime. The code can be supplied in a programming language that can be selected to enable the code to execute in a precompiled or as-compiled fashion.
[0142] Embodiments of the systems and methods provided herein, such as the computer system 701, can be embodied in programming. Various aspects of the technology may be thought of as “products” or “articles of manufacture” typically in the form of machine (or processor) executable code and / or associated data that is carried on or embodied in a type of machine readable medium. Machine-executable code can be stored on an electronic storage unit, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. “Storage” type media can include any or all of the tangible memory of the computers, processors or the like, or associated modules thereof, such as various semiconductor memories, tape drives, or disk drives, which may provide non-transitory storage at any time for the software programming. All or portions of the software may at times be communicated through the Internet or various other telecommunication networks. Such communications, for example, may enable loading of the software from one computer or processor into another, for example, from a management server or host computer into the computer platform of an application server. Thus, another type of media that may bear the software elements includes optical, electrical andelectromagnetic waves, such as used across physical interfaces between local devices, through wired and optical landline networks and over various air-links. The physical elements that carry such waves, such as wired or wireless links, optical links or the like, also may be considered as media bearing the software. As used herein, unless restricted to non-transitory, tangible “storage” media, terms such as computer or machine “readable medium” refer to any medium that participates in providing instructions to a processor for execution.
[0143] Hence, a machine readable medium, such as computer-executable code, may take many forms, including a tangible storage medium, a carrier wave medium or physical transmission medium. Non-volatile storage media include, for example, optical or magnetic disks, such as any of the storage devices in any computer(s) or the like, such as may be used to implement the databases, etc. shown in the drawings. Volatile storage media include dynamic memory, such as main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that comprise a bus within a computer system. Carrier-wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media therefore include for example: a floppy disk, a flexible disk, hard disk, magnetic tape, any other magnetic medium, a CD-ROM, DVD or DVD-ROM, any other optical medium, punch cards paper tape, any other physical storage medium with patterns of holes, a RAM, a ROM, a PROM and EPROM, a FLASH-EPROM, any other memory chip or cartridge, a carrier wave transporting data or instructions, cables or links transporting such a carrier wave, or any other medium from which a computer may read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0144] The computer system 701 can include or be in communication with an electronic display 735 that comprises a user interface (UI) 740 for providing, for example, a visual display indicative of the cancer status a subject based on small RNA sequencing reads derived from a sample from the subject, a visual display indicative of a planned course of treatment for a subject based on small RNA sequencing reads derived from a sample from the subject, a visual display indicative of monitoring the cancer status of a subject over time (e.g., before and after treatment) based on small RNA sequencing reads derived from a sample from the subject, a visual display indicated of a measurement the efficacy or tolerability of a therapeutic intervention for an individual based on small RNA sequencing reads derived from a sample from the subject, or a visual display indicative of processing a biological sample asdescribed herein. Examples of UIs include, without limitation, a graphical user interface (GUI) and web-based user interface.
[0145] Methods and systems of the present disclosure can be implemented by way of one or more algorithms. An algorithm can be implemented by way of software upon execution by the central processing unit 705. The algorithm can, for example, determine the cancer status a subject from small RNA expression data derived from a sample from the subject, plan a course of treatment for a subject based on small RNA sequencing reads derived from a sample from the subject, monitor the cancer status of a subject over time (e.g., before and after treatment) based on small RNA sequencing reads derived from a sample from the subject, measure the efficacy or tolerability of a therapeutic intervention for an individual based on small RNA sequencing reads derived from a sample from the subject, or process a biological sample as described herein.
[0146] Table 2 lists 10,334 non-coding RNAs associated with cancer v. non-cancer status in a cohort of cancer and non-cancer samples, either through frequency comparisons (binary presence or absence) or differential expression. In some cases, the non-coding RNAs listed in Table 2 are associated with cancer v. non-cancer with a false discovery rate below 0.1. Whether a subject has residual cancer or minimal residual disease can be determined based on the presence or level of one or more non-coding RNAs from Table 2 detected in a sample from the subject.Table 2: non-coding RNAs associated with cancer v. non-cancer status.-Ill-
Claims
CLAIMSWHAT IS CLAIMED IS:
1. A method, comprising:(a) providing a sample from a subject, wherein said subject has received treatment for cancer and has or is suspected of having residual cancer cells of said cancer, and wherein said sample comprises one or more non-coding RNA molecules that are present in said residual cancer cells and have a 90th percentile expression in non-cancerous cells of below 0.5 count per-million reads (cpm);(b) sequencing said one or more non-coding RNA molecules, thereby generating a data set comprising data corresponding to a presence of said one or more noncoding RNA molecules in the sample;(c) inputting said data set into a trained algorithm to generate a classification that includes an indication that said sample is positive or negative for said residual cancer cells; and(d) electronically outputting said classification in a report.
2. The method of claim 1, wherein said classification further comprises a prediction of a likelihood that the residual cancer cells will proliferate.
3. The method of claim 1, wherein said classification further comprises a prediction of survival of said subject.
4. The method of claim 1, wherein said classification further comprises a prediction of a clinical characteristic comprising a cancer stage, a tumor (T) stage, a node invasion score, a metastasis score, a tumor grade score, a tumor size score, a cancer site score, or any combination thereof.
5. The method of claim 1, wherein said classification further comprises a prediction of a pathologic complete response (pCR).
6. The method of claim 1, wherein said treatment comprises a neoadjuvant chemotherapy.
7. The method of claim 1, wherein said one or more non-coding RNA molecules have a length of less than 200 nucleotides.
8. The method of claim 1, wherein said one or more non-coding RNA molecules have a length of between 50 and 100 nucleotides.
9. The method of claim 1, wherein said sequencing in (b) comprises subjecting said one or more non-coding RNA molecules to reverse transcription to generate one or more complementary deoxyribonucleic acid (DNA) molecules.
10. The method of claim 9, wherein said sequencing in (b) further comprises sequencing said one or more cDNA molecules or derivatives thereof.
11. The method of claim 10, further comprising, after (b), amplifying a cDNA molecule of said one or more cDNA molecules.
12. The method of claim 11, wherein said amplifying comprises performing a polymerase chain reaction (PCR).
13. The method of claim 12, wherein said amplifying comprises rolling circle amplification.
14. The method of claim 1, wherein said sequencing in (b) further comprises sequencing-by-synthesis.
15. The method of claim 1, wherein said sample is a cell-free sample.
16. The method of claim 1, wherein said sample comprises serum.
17. The method of claim 1, wherein said sample comprises plasma.
18. The method of claim 1, wherein said sample comprises urine.
19. The method of claim 1, wherein said sample comprises lymph.
20. The method of claim 1, wherein said sample comprises saliva.
21. The method of claim 1, wherein a total volume of said sample is between 20 microliters and 2 milliliters.
22. The method of claim 1, wherein said sample comprises plasma, and wherein a total volume of said sample is between 100 microliters and 1 milliliter.
23. The method of claim 1, wherein said trained algorithm comprises a machine learning model.
24. The method of claim 23, wherein said machine learning model comprises a neural network.
25. The method of claim 24, wherein said neural network comprises a regressor neural network, a multi-head neural network, or any combination thereof.
26. The method of claim 1, wherein said trained algorithm is trained using noncoding RNA derived from a treatment-naive patient sample.
27. The method of claim 26, wherein said treatment-naive patient sample comprises serum.
28. The method of claim 1, wherein said trained algorithm is trained across multiple demographics, cancer types, and cancer stages.
29. The method of claim 1, wherein said residual cancer cells comprise breast cancer cells.
30. The method of claim 29, wherein said breast cancer cells comprise triple-negative breast cancer cells.
31. The method of claim 1, further comprising (e) providing said report to a user.
32. The method of claim 31, wherein said user is a clinician.
33. The method of claim 31, wherein said providing further comprises delivering said report to a computer interface of said user.
34. The method of claim 1, wherein said trained algorithm is trained on clinical characteristics comprising a cancer stage, a T stage, a node invasion score, a metastasis score, a tumor grade score, a tumor size score, a cancer site score, or any combination thereof.
35. A method, comprising:(a) providing a sample from a subject, wherein said subject has received treatment for cancer and has or is suspected of having residual cancer cells of said cancer, wherein said sample comprises one or more non-coding RNA molecules;(b) sequencing said sample, thereby obtaining a first data set comprising a first set of one or more non-coding RNA sequences of said one or more non-coding RNA molecules in said sample;(c) providing a second data set comprising a second set of one or more noncoding RNA sequences of one or more non-coding RNA molecules that are (i) present in residual cancer cells and (ii) have a 90th percentile expression in normal samples below 0.5 count-per-million (cpm) reads;(d) providing said first data set and said second data set into a trained algorithm to generate a classification that includes an indication that said sample is positive or negative for said residual cancer cells; and(e) electronically outputting said classification in a report.
36. The method of claim 35, wherein said classification further comprises a prediction of a likelihood that the residual cancer cells will proliferate.
37. The method of claim 35, wherein said classification further comprises a prediction of survival of said subject.
38. The method of claim 35, wherein said classification further comprises a prediction of a clinical characteristic comprising a cancer stage, a tumor (T) stage, a node invasion score, a metastasis score, a tumor grade score, a tumor size score, a cancer site score, or any combination thereof.
39. The method of claim 35, wherein said classification further comprises a prediction of a pathologic complete response (pCR).
40. The method of claim 35, wherein said treatment comprises a neoadjuvant chemotherapy.
41. The method of claim 35, wherein said one or more non-coding RNA molecules have a length of less than 200 nucleotides.
42. The method of claim 35, wherein said one or more non-coding RNA molecules have a length of between 50 and 100 nucleotides.
43. The method of claim 35, wherein said sequencing in (b) comprises subjecting said one or more non-coding RNA molecules to reverse transcription to generate one or more complementary deoxyribonucleic acid (DNA) molecules.
44. The method of claim 43, wherein said sequencing in (b) further comprises sequencing said one or more cDNA molecules or derivatives thereof.
45. The method of claim 44, further comprising, after (b), amplifying a cDNA molecule of said one or more cDNA molecules.
46. The method of claim 45, wherein said amplifying comprises performing a polymerase chain reaction (PCR).
47. The method of claim 46, wherein said amplifying comprises rolling circle amplification.
48. The method of claim 35, wherein said sequencing in (b) further comprises sequencing-by-synthesis.
49. The method of claim 35, wherein said sample is a cell-free sample.
50. The method of claim 35, wherein said sample comprises serum.
51. The method of claim 35, wherein said sample comprises plasma.
52. The method of claim 35, wherein said sample comprises urine.
53. The method of claim 35, wherein said sample comprises lymph.
54. The method of claim 35, wherein said sample comprises saliva.
55. The method of claim 35, wherein a total volume of said sample is between 20 microliters and 2 milliliters.
56. The method of claim 35, wherein said sample comprises plasma, and wherein a total volume of said sample is between 100 microliters and 1 milliliter.
57. The method of claim 35, wherein said trained algorithm comprises a machine learning model.
58. The method of claim 57, wherein said machine learning model comprises a neural network.
59. The method of claim 58, wherein said neural network comprises a regressor neural network, a multi-head neural network, or any combination thereof.
60. The method of claim 35, wherein said trained algorithm is trained using noncoding RNA derived from a treatment-naive patient sample.
61. The method of claim 60, wherein said treatment-naive patient sample comprises serum.
62. The method of claim 35, wherein said trained algorithm is trained across multiple demographics, cancer types, and cancer stages.
63. The method of claim 35, wherein said residual cancer cells comprise breast cancer cells.
64. The method of claim 63, wherein said breast cancer cells comprise triple-negative breast cancer cells.
65. The method of claim 35, further comprising (f) providing said report to a user.
66. The method of claim 65, wherein said user is a clinician.
67. The method of claim 65, wherein said providing further comprises delivering said report to a computer interface of said user.
68. The method of claim 35, wherein said trained algorithm is trained on clinical characteristics comprising a cancer stage, a tumor (T) stage, a node invasion score, a metastasis score, a tumor grade score, a tumor size score, a cancer site score, or any combination thereof.
69. A method, comprising:(a) providing a sample from a subject, wherein said subject has received treatment for cancer and has or is suspected of having residual cancer cells of said cancer, and wherein said sample comprises one or more non-coding RNA molecules having a sequence of any one of SEQ ID NOs: 1-10,334;(b) sequencing said one or more non-coding RNA molecules, thereby generating a data set comprising data corresponding to a presence of said one or more noncoding RNA molecules in the sample;(c) inputting said data set into a trained algorithm to generate a classification that includes an indication that said sample is positive or negative for said residual cancer cells; and(d) electronically outputting said classification in a report.
70. The method of claim 69, wherein said classification further comprises a prediction of a likelihood that the residual cancer cells will proliferate.
71. The method of claim 69, wherein said classification further comprises a prediction of survival of said subject.
72. The method of claim 69, wherein said classification further comprises a prediction of a clinical characteristic comprising a cancer stage, a tumor (T) stage, a nodeinvasion score, a metastasis score, a tumor grade score, a tumor size score, a cancer site score, or any combination thereof.
73. The method of claim 69, wherein said classification further comprises a prediction of a pathologic complete response (pCR).
74. The method of claim 69, wherein said treatment comprises a neoadjuvant chemotherapy.
75. The method of claim 69, wherein said one or more non-coding RNA molecules have a length of less than 200 nucleotides.
76. The method of claim 69, wherein said one or more non-coding RNA molecules have a length of between 50 and 100 nucleotides.
77. The method of claim 69, wherein said sequencing in (b) comprises subjecting said one or more non-coding RNA molecules to reverse transcription to generate one or more complementary deoxyribonucleic acid (DNA) molecules.
78. The method of claim 77, wherein said sequencing in (b) further comprises sequencing said one or more cDNA molecules or derivatives thereof.
79. The method of claim 78, further comprising, after (b), amplifying a cDNA molecule of said one or more cDNA molecules.
80. The method of claim 79, wherein said amplifying comprises performing a polymerase chain reaction (PCR).
81. The method of claim 80, wherein said amplifying comprises rolling circle amplification.
82. The method of claim 69, wherein said sequencing in (b) further comprises sequencing-by-synthesis.
83. The method of claim 69, wherein said sample is a cell-free sample.
84. The method of claim 69, wherein said sample comprises serum.
85. The method of claim 69, wherein said sample comprises plasma.
86. The method of claim 69, wherein said sample comprises urine.
87. The method of claim 69, wherein said sample comprises lymph.
88. The method of claim 69, wherein said sample comprises saliva.
89. The method of claim 69, wherein a total volume of said sample is between 20 microliters and 2 milliliters.
90. The method of claim 69, wherein said sample comprises plasma, and wherein a total volume of said sample is between 100 microliters and 1 milliliter.
91. The method of claim 69, wherein said trained algorithm comprises a machine learning model.
92. The method of claim 91, wherein said machine learning model comprises a neural network.
93. The method of claim 92, wherein said neural network comprises a regressor neural network, a multi-head neural network, or any combination thereof.
94. The method of claim 69, wherein said trained algorithm is trained using noncoding RNA derived from a treatment-naive patient sample.
95. The method of claim 94, wherein said treatment-naive patient sample comprises serum.
96. The method of claim 69, wherein said trained algorithm is trained across multiple demographics, cancer types, and cancer stages.
97. The method of claim 69, wherein said residual cancer cells comprise breast cancer cells.
98. The method of claim 97, wherein said breast cancer cells comprise triple-negative breast cancer cells.
99. The method of claim 69, further comprising (e) providing said report to a user.
100. The method of claim 99, wherein said user is a clinician.
101. The method of claim 99, wherein said providing further comprises delivering said report to a computer interface of said user.
102. The method of claim 69, wherein said trained algorithm is trained on clinical characteristics comprising a cancer stage, a T stage, a node invasion score, a metastasis score, a tumor grade score, a tumor size score, a cancer site score, or any combination thereof.
103. A method, comprising:(a) providing a sample from a subject, wherein said subject has received treatment for cancer and has or is suspected of having residual cancer cells of said cancer, wherein said sample comprises one or more non-coding RNA molecules;(b) sequencing said sample, thereby obtaining a first data set comprising a first set of one or more non-coding RNA sequences of said one or more non-coding RNA molecules in said sample;(c) providing a second data set comprising a second set of one or more noncoding RNA sequences of one or more non-coding RNA molecules having a sequence of any one of SEQ ID NOs: 1-10,334;(d) providing said first data set and said second data set into a trained algorithm to generate a classification that includes an indication that said sample is positive or negative for said residual cancer cells; and(e) electronically outputting said classification in a report.
104. The method of claim 103, wherein said classification further comprises a prediction of a likelihood that the residual cancer cells will proliferate.
105. The method of claim 103, wherein said classification further comprises a prediction of survival of said subject.
106. The method of claim 103, wherein said classification further comprises a prediction of a clinical characteristic comprising a cancer stage, a tumor (T) stage, a node invasion score, a metastasis score, a tumor grade score, a tumor size score, a cancer site score, or any combination thereof.
107. The method of claim 103, wherein said classification further comprises a prediction of a pathologic complete response (pCR).
108. The method of claim 103, wherein said treatment comprises a neoadjuvant chemotherapy.
109. The method of claim 103, wherein said one or more non-coding RNA molecules have a length of less than 200 nucleotides.
110. The method of claim 103, wherein said one or more non-coding RNA molecules have a length of between 50 and 100 nucleotides.
111. The method of claim 103, wherein said sequencing in (b) comprises subjecting said one or more non-coding RNA molecules to reverse transcription to generate one or more complementary deoxyribonucleic acid (DNA) molecules.
112. The method of claim 111, wherein said sequencing in (b) further comprises sequencing said one or more cDNA molecules or derivatives thereof.
113. The method of claim 112, further comprising, after (b), amplifying a cDNA molecule of said one or more cDNA molecules.
114. The method of claim 113, wherein said amplifying comprises performing a polymerase chain reaction (PCR).
115. The method of claim 114, wherein said amplifying comprises rolling circle amplification.
116. The method of claim 103, wherein said sequencing in (b) further comprises sequencing-by-synthesis.
117. The method of claim 103, wherein said sample is a cell -free sample.
118. The method of claim 103, wherein said sample comprises serum.
119. The method of claim 103, wherein said sample comprises plasma.
120. The method of claim 103, wherein said sample comprises urine.
121. The method of claim 103, wherein said sample comprises lymph.
122. The method of claim 103, wherein said sample comprises saliva.
123. The method of claim 103, wherein a total volume of said sample is between 20 microliters and 2 milliliters.
124. The method of claim 103, wherein said sample comprises plasma, and wherein a total volume of said sample is between 100 microliters and 1 milliliter.
125. The method of claim 103, wherein said trained algorithm comprises a machine learning model.
126. The method of claim 125, wherein said machine learning model comprises a neural network.
127. The method of claim 126, wherein said neural network comprises a regressor neural network, a multi-head neural network, or any combination thereof.
128. The method of claim 103, wherein said trained algorithm is trained using noncoding RNA derived from a treatment-naive patient sample.
129. The method of claim 128, wherein said treatment-naive patient sample comprises serum.
130. The method of claim 103, wherein said trained algorithm is trained across multiple demographics, cancer types, and cancer stages.
131. The method of claim 103, wherein said residual cancer cells comprise breast cancer cells.
132. The method of claim 131, wherein said breast cancer cells comprise triplenegative breast cancer cells.
133. The method of claim 103, further comprising (f) providing said report to a user.
134. The method of claim 133, wherein said user is a clinician.
135. The method of claim 133, wherein said providing further comprises delivering said report to a computer interface of said user.
136. The method of claim 103, wherein said trained algorithm is trained on clinical characteristics comprising a cancer stage, a tumor (T) stage, a node invasion score, a metastasis score, a tumor grade score, a tumor size score, a cancer site score, or any combination thereof.