Methods and systems for patient stratification

WO2026190756A1PCT designated stage Publication Date: 2026-09-17PACIFIC EDGE LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2026/052492
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-14
Filing Date
2026-03-13
Publication Date
2026-09-17

Smart Images

  • Figure IB2026052492_17092026_PF_FP_ABST
    Figure IB2026052492_17092026_PF_FP_ABST
Patent Text Reader

Abstract

Provided are methods, systems, and kits for analyzing genetic data and transcriptomic data derived from a sample of a subject with a trained machine learning algorithm to determine a degree of bladder cancer risk in the subject.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No. 70528-701.601METHODS AND SYSTEMS FOR PATIENT STRATIFICATION CROSS-REFERENCE

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 772,292, filed March 14, 2025, which application is incorporated herein by reference.BACKGROUND

[0002] Cystoscopy is a common diagnostic test for bladder cancers, including urothelial carcinoma (UC). Cystoscopy is an invasive procedure involving inserting a thin, flexible tube called a cystoscope into the urethra and advanced into the bladder. The cystoscope has a light and a camera that allows visualization inside of the bladder. Cystoscopy poses many risks to a patient, including a risk of infection and / or bleeding, and causes a significant amount of discomfort or pain during the procedure. Although imaging may be used to visualize large tumors in the bladder and surrounding tissues, imaging such as ultrasound, cannot definitively diagnose bladder cancer, such as UC.

[0003] Patients presenting with hematuria (e.g., blood in the urine) may be subject to further testing under the American Urological Association (AU A) guidelines, like cystoscopy and pathological confirmation, to confirm or rule out a diagnosis of bladder cancer. Pathological confirmation includes a biopsy of the bladder of the subject, which also poses several risks to patients. Moreover, a high proportion of patients with hematuria who undergo this further testing have inconclusive findings, such as atypical cytology or equivocal cystoscopy, requiring further invasive tests (e.g., biopsy).SUMMARY

[0004] Aspects disclosed herein provide methods for analyzing a sample, the methods comprising: a) assaying the sample or a first portion of the sample obtained from a subject that has bladder cancer or is suspected of having bladder cancer with a genotyping assay to detect one or more genotypes at one or more polymorphisms, thereby generating a first data set, wherein the one or more polymorphisms comprise rsl21913482, rsl21913483, rsl21913479, rsl21913485, rsl242535815, rsl561215364, or a polymorphism in linkage disequilibrium therewith as determined by a coefficient of determination R2 of at least 0.85, or any combination thereof; b) assaying the sample or a second portion of the sample with a transcriptomic assay to determine a quantitative measure of one or more transcriptomic markers, thereby generating a second data set, wherein the one or more transcriptomic markers comprise Midkine (MDK),Attomey Docket No. 70528-701.601Cyclin Dependent Kinase 1 (CDK1), Insulin Like Growth Factor Binding Protein 5 (IGFBP5), Homeobox Al 3 (HOXA13), or C-X-C Motif Chemokine Receptor 2 (CXCR2), or any combination thereof; and c) analyzing the first data set and the second data set using a trained machine learning algorithm to determine an output indicative of the subject as having a degree of cancer risk. In some embodiments, the method generating an electronic report comprising the output indicative of the subject as having the degree of cancer risk. In some embodiments, the subject has or is suspected of having hematuria. In some embodiments, the one or more genotypes are heterozygous for a risk allele at the one or more polymorphisms. In some embodiments, the one or more polymorphisms comprise two or more of rsl21913482, rsl21913483, rsl21913479, rsl21913485, rsl242535815, rsl561215364, or a polymorphism in linkage disequilibrium therewith as determined by a coefficient of determination R2 of at least 0.85, or any combination thereof. In some embodiments, the one or more polymorphisms comprise rsl21913482, rsl21913483, rsl21913479, rsl21913485, rsl242535815, and rsl561215364. In some embodiments, the one or more transcriptomic markers comprise MDK, CDK1, IGFBP5, HOXA13, and CXCR2. In some embodiments, the one or more transcriptomic markers comprise MDK, CDK1, IGFBP5, HOXA13, and CXCR2; and the one or more polymorphisms comprise rsl21913482, rsl21913483, rsl21913479, rsl21913485, rsl242535815, and rsl561215364. In some embodiments, the quantitative measure comprises a level or an amount of the one or more transcriptomic markers. In some embodiments, the trained machine learning algorithm comprises a trained machine learning classifier, wherein the degree of cancer risk comprises a categorical cancer risk selected from among a plurality of distinct categorical cancer risks, and wherein determining the output comprises assigning a classification of the subject as having the categorical cancer risk selected from among the plurality of distinct categorical cancer risks. In some embodiments, the plurality of distinct categorical cancer risks comprises a high cancer risk, an intermediate cancer risk, or a low cancer risk, or any combination thereof. In some embodiments, the plurality of distinct categorical cancer risks are determined with reference to thresholds based at least in part on a highest calculated Youden index of a Receiver Operating Characteristic (ROC) curve for the trained machine learning algorithm. In some embodiments, the assigned classification is the high cancer risk or the intermediate cancer risk, and the method further comprises administering a treatment to the subject capable of treating the bladder cancer. In some embodiments, the treatment comprises a surgical resection, a chemotherapy, an immunotherapy, a radiation therapy, a targeted therapy, or a combination thereof. In some embodiments, the classification has a specificity of at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, or at least about 95%. In some embodiments, the classification has a sensitivity of at least about 85%,Attomey Docket No. 70528-701.601at least about 90%, or at least about 95%. In some embodiments, the classification has a negative predictive value (NPV) of at least about 95%. In some embodiments, the classification has a negative predictive value (NPV) of at least about 96.5%. In some embodiments, the classification a positive predictive value (PPV) of at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, or at least about 95%. In some embodiments, the trained machine learning algorithm comprises a non-parametric model. In some embodiments, the trained machine learning algorithm comprises a boosted regression tree (BRT), a Bayesian Additive Regression Tree (BART), a random forest (RF) algorithm, a boosted tree algorithm, a gradient boosting machine (GBM), a Bayesian Model Averaging (BMA) with a decision tree, a classification and regression tree (CART) model, a support vector machine (SVM), a Linear Discriminate Analysis (LDA), a Logistic Regression (LogReg), a K-nearest 5 neighbors (Kn5n), a partition tree classifier (TREE), or a partially fixed BART, or any combination thereof. In some embodiments, the trained machine learning algorithm comprises the BART, the RF, the GBM, or a Bayesian CART. In some embodiments, the analyzing the first data set and the second data set using the trained machine learning algorithm is performed with a single trained machine learning algorithm. In some embodiments, the trained machine learning algorithm is an integrated algorithm or an ensemble machine learning algorithm comprising two or more algorithms. In some embodiments, the trained machine learning algorithm comprises an inverse logit function. In some embodiments, the inverse logit function is an inverse logit of a second-order polynomial. In some embodiments, the second-order polynomial comprises coefficients obtained by fitting a logistic regression model. In some embodiments, the inverse logit function determines an estimated probability of the subject having cancer. In some embodiments, the method further comprises risk stratifying the subject based at least in part on thresholding the estimated probability of the subject having cancer. In some embodiments, the degree of cancer risk comprises an estimated probability of cancer. In some embodiments, the sample is a urine sample. In some embodiments, the method further comprises collecting the urine sample from the subject. In some embodiments, the transcriptomic assay comprises a droplet digital polymerase chain reaction (PCR). In some embodiments, the genotyping assay comprises a quantitative reverse transcription analysis. In some embodiments, the bladder cancer is urothelial bladder cancer. In some embodiments, the method further comprises training the trained machine learning algorithm prior to (a). In some embodiments, the training comprises: receiving a genotype data and transcriptomic data from a plurality of training samples comprising case samples obtained from subjects with the bladderAttomey Docket No. 70528-701.601cancer, and control samples obtained from subjects without the bladder cancer; creating a first training set comprising covariate data corresponding to the genotype data and the transcriptomic data; and training a machine learning algorithm with the covariate data to determine an output indicative of the subject as having a degree of cancer risk. In some embodiments, the trained machine learning algorithm is trained with a training set that is independent of the sample. In some embodiments, the trained machine learning algorithm has a performance characterized by a receiver operating characteristic (ROC) curve having an average or median area under the curve (AUC) of at least 0.85. In some embodiments, the trained machine learning algorithm has a performance characteristic comprising a specificity of at least about 70% when the trained machine learning algorithm is trained on bladder cancer samples and non-cancer samples. In some embodiments, the trained machine learning algorithm has a performance characteristic comprising a sensitivity of at least about 85% when the trained machine learning algorithm is trained on bladder cancer samples and non-cancer samples. In some embodiments, the trained machine learning algorithm has a performance characteristic comprising a negative predictive value (NPV) of at least about 95% when the trained machine learning algorithm is trained on bladder cancer samples and non-cancer samples. In some embodiments, the trained machine learning algorithm has a performance characteristic comprising a negative predictive value (NPV) of at least about 96.5% when the trained machine learning algorithm is trained on bladder cancer samples and non-cancer samples. In some embodiments, the trained machine learning algorithm has a performance characteristic comprising a positive predictive value (PPV) of at least about 15% when the trained machine learning algorithm is trained on bladder cancer samples and non-cancer samples. In some embodiments, the method further comprises identifying the subject as not having the bladder cancer or having a low risk of developing the bladder cancer, based at least in part on the assigned classification being the low cancer risk, thereby avoiding having to perform a cystoscopy on the subject. In some embodiments, the method further comprises performing a cystoscopy on the subject to confirm a diagnosis of the bladder cancer, based at least in part on the assigned classification being the high cancer risk. In some embodiments, the method further comprises calculating a risk score corresponding to the high cancer risk, wherein the risk score exceeds the threshold of 0.54. In some embodiments, the method further comprises calculating a risk score corresponding to the intermediate cancer risk, wherein the risk score is from 0.15 to 0.54. In some embodiments, the method further comprises further comprising calculating a risk score corresponding to the low cancer risk, wherein the risk score is below 0.15.

[0005] Aspects disclosed herein provide systems comprising: at least one processor, an operating system configured to perform executable instructions, a memories storing machine-Attomey Docket No. 70528-701.601executable code that, when executed, causes the at least one processor to: a) receive a first data set obtained from assaying a sample or a first portion of the sample obtained from a subject that has bladder cancer or is suspected of having bladder cancer with a genotyping assay to detect one or more genotypes at one or more polymorphisms, wherein the one or more polymorphisms comprise rsl21913482, rsl21913483, rsl21913479, rsl21913485, rsl242535815,rs 1561215364, or a polymorphism in linkage disequilibrium therewith as determined by a coefficient of determination R2 of at least 0.85, or any combination thereof; b) receive a second data set obtained from assaying the sample or a second portion of the sample obtained from the subject that has bladder cancer or is suspected of having bladder cancer with a transcriptomic assay to determine a quantitative measure of one or more transcriptomic markers to generate the second data set, wherein the one or more transcriptomic markers comprise Midkine (MDK), Cyclin Dependent Kinase 1 (CDK1), Insulin Like Growth Factor Binding Protein 5 (IGFBP5), Homeobox Al 3 (HOXA13), or C-X-C Motif Chemokine Receptor 2 (CXCR2), or any combination thereof; and c) analyze the first data set and the second data set using a trained machine learning algorithm to determine an output indicative of the subject as having a degree of cancer risk. In some embodiments, the one or more genotypes are heterozygous for a risk allele at the one or more polymorphisms. In some embodiments, the one or more polymorphisms comprise two or more of rsl21913482, rsl21913483, rsl21913479, rsl21913485, rsl242535815, rsl561215364, or a polymorphism in linkage disequilibrium therewith as determined by a coefficient of determination R2of at least 0.85, or any combination thereof. In some embodiments, the one or more polymorphisms comprise rsl21913482, rsl21913483, rsl21913479, rsl21913485, rsl242535815, and rsl561215364. In some embodiments, the one or more transcriptomic markers comprise MDK, CDK1, IGFBP5, HOXA13, and CXCR2. In some embodiments, the one or more transcriptomic markers comprises MDK, CDK1, IGFBP5, HOXA13, and CXCR2; and the one or more polymorphisms comprises rsl21913482, rsl21913483, rsl21913479, rsl21913485, rsl242535815, and rsl561215364. In some embodiments, the quantitative measure comprises a level or an amount of the one or more transcriptomic markers. In some embodiments, wherein the trained machine learning algorithm comprises a trained machine learning classifier, wherein the degree of cancer risk comprises a categorical cancer risk selected from among a plurality of distinct categorical cancer risks, and wherein determining the output comprises assigning a classification of the subject as having the categorical cancer risk selected from among the plurality of distinct categorical cancer risks. In some embodiments, the plurality of distinct categorical cancer risks comprises a high cancer risk, an intermediate cancer risk, or a low cancer risk, or any combination thereof. In some embodiments, the plurality of distinct categorical cancer risks are determined with reference toAttomey Docket No. 70528-701.601thresholds based at least in part on a highest calculated Youden index of a Receiver Operating Characteristic (ROC) curve for the trained machine learning algorithm. In some embodiments, the assigned classification is the high cancer risk or the intermediate cancer risk, and the machine-executable code, when executed, causes the at least one processor to output a report comprising a recommended treatment of the bladder cancer. In some embodiments, the treatment comprises a surgical resection, a chemotherapy, an immunotherapy, a radiation therapy, a targeted therapy, or a combination thereof. In some embodiments, the trained machine learning algorithm has a performance characterized by a receiver operating characteristic (ROC) curve having an average or median area under the curve (AUC) of at least about 0.85. In some embodiments, the trained machine learning algorithm has a performance characterized by a specificity of at least about 70%. In some embodiments, the sample is classified as a high cancer risk, an intermediate cancer risk, or a low cancer risk at a sensitivity of at least about 85%. In some embodiments, the sample is classified as a high cancer risk, an intermediate cancer risk, or a low cancer risk at a negative predictive value (NPV) of at least about 96.5%. In some embodiments, the sample is classified as a high cancer risk, an intermediate cancer risk, or a low cancer risk at a positive predictive value (PPV) of at least about 15%. In some embodiments, the subject has or is suspected of having hematuria. In some embodiments, the trained machine learning algorithm comprises a non-parametric model. In some embodiments, the trained machine learning algorithm comprises a boosted regression tree (BRT), a Bayesian Additive Regression Tree (BART), a random forest (RF) algorithm, a boosted tree algorithm, a gradient boosting machine (GBM), a Bayesian Model Averaging (BMA) with a decision tree, a classification and regression tree (CART) model, a support vector machine (SVM), a Linear Discriminate Analysis (LDA), a Logistic Regression (LogReg), a K-nearest 5 neighbors (Kn5n), a partition tree classifier (TREE), or a partially fixed BART, or any combination thereof. In some embodiments, the trained machine learning algorithm comprises the BART, the RF, the GBM, or a Bayesian CART. In some embodiments, the analyzing the first data set and the second data set using the trained machine learning algorithm is performed with a single trained machine learning algorithm. In some embodiments, the trained machine learning algorithm is an integrated algorithm comprising two or more algorithms. In some embodiments, the trained machine learning algorithm comprises an inverse logit function. In some embodiments, the inverse logit function is an inverse logit of a second-order polynomial. In some embodiments, the second-order polynomial comprises coefficients obtained by fitting a logistic regression model. In some embodiments, the inverse logit function determines an estimated probability of the subject having cancer. In some embodiments, the machine-executable code, when executed, causes the at least one processor to risk stratifying the subject based at least in part onAttomey Docket No. 70528-701.601thresholding the estimated probability of the subject having cancer. In some embodiments, the degree of cancer risk comprises an estimated probability of cancer. In some embodiments, the sample is a urine sample. In some embodiments, the transcriptomic assay comprises a droplet digital polymerase chain reaction (PCR). In some embodiments, the genotyping assay comprises a quantitative reverse transcription analysis. In some embodiments, the bladder cancer is urothelial bladder cancer. In some embodiments, the trained machine learning algorithm comprises: a genotype data and transcriptomic data from a plurality of training samples comprising case samples obtained from subjects with the bladder cancer, and control samples obtained from subjects without the bladder cancer; and a first training set comprising covariate data corresponding to the genotype data and the transcriptomic data, wherein the machine learning algorithm is configured to be trained by covariate data to determine an output indicative of the subject as having a degree of cancer risk. In some embodiments, the trained machine learning algorithm is trained with a training set that is independent of the sample. In some embodiments, the trained machine learning algorithm has a performance characterized by a receiver operating characteristic (ROC) curve having an average or median area under the curve (AUC) of at least about 0.85. In some embodiments, the trained machine learning algorithm has a performance characteristic comprising a specificity of at least about 70% when the trained machine learning algorithm is trained on bladder cancer samples and non-cancer samples. In some embodiments, the trained machine learning algorithm has a performance characteristic comprising a sensitivity of at least about 85% when the trained machine learning algorithm is trained on bladder cancer samples and non-cancer samples. In some embodiments, the trained machine learning algorithm has a performance characteristic comprising a negative predictive value (NPV) of at least about 95% when the trained machine learning algorithm is trained on bladder cancer samples and non-cancer samples. In some embodiments, the trained machine learning algorithm has a performance characteristic comprising a negative predictive value (NPV) of at least about 96.5% when the trained machine learning algorithm is trained on bladder cancer samples and non-cancer samples. In some embodiments, the trained machine learning algorithm has a performance characteristic comprising a positive predictive value (PPV) of at least 15% when the trained machine learning algorithm is trained on bladder cancer samples and non-cancer samples. In some embodiments, the classification is configured to identify the subject as not having the bladder cancer or having a low risk of developing the bladder cancer based, at least in part, on the assigned classification being the low cancer risk, thereby avoiding having to perform a cystoscopy on the subject. In some embodiments, the system further comprising a cystoscope that is configured to be used in a cystoscopy on the subject to confirm a diagnosis of the bladder cancer, based at least in part on the assigned classification being theAttomey Docket No. 70528-701.601high or intermediate cancer risk. In some embodiments, the machine-executable code, when executed, causes the at least one processor to output a report comprising a risk score corresponding to the high cancer risk, wherein the risk score exceeds the threshold of 0.54. In some embodiments, the machine-executable code, when executed, causes the at least one processor to output a report comprising a risk score corresponding to the intermediate cancer risk, wherein the risk score is from 0.15 to 0.54. In some embodiments, the machine-executable code, when executed, causes the at least one processor to output a report comprising a risk score corresponding to the low cancer risk, wherein the risk score is below 0.15.INCORPORATION BY REFERENCE

[0006] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.BRIEF DESCRIPTION OF THE DRAWINGS

[0007] The novel features of the inventive concepts are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present inventive concepts will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the inventive concepts are utilized, and the accompanying drawings of which:

[0008] FIG. 1 shows a non-limiting example of a computing device; in this case, a device with one or more processors, memory, storage, and a network interface.

[0009] FIG. 2 shows a non-limiting example of a web / mobile application provision system; in this case, a system providing browser-based and / or native mobile user interfaces; and

[0010] FIG. 3 shows a non-limiting example of a cloud-based web / mobile application provision system; in this case, a system comprising an elastically load balanced, auto-scaling web server and application server resources as well synchronously replicated databases.

[0011] FIG. 4 shows an example of a decision tree for a clinician in analyzing the risk of a subject’s cancer.

[0012] FIG. 5 shows an example of a linearity graph of the measured versus known DNA concentrations for the six single nucleotide polymorphisms (SNPs).Attomey Docket No. 70528-701.601

[0013] FIG. 6 shows an example of a flow diagram of patient selection for the analysis set.

[0014] FIG. 7 shows an example of a calibration curve of pre-specified thresholds.

[0015] FIG. 8 shows an example of receiver operating characteristic (ROC) curves.

[0016] FIG. 9 shows an example of a multi-modal diagnostic workflow compared with AUA MH risk criteria or GH status and hematuria status risk stratification.DETAILED DESCRIPTION

[0017] Provided herein are methods, systems, and kits for analyzing genetic data and transcriptomic data derived from a sample of a subject with a trained machine learning algorithm to determine a degree of cancer risk in the subject. The subject may be suspected of having bladder cancer, which may include urothelial cancer (UC). The subject may be a patient, who has hematuria. Hematuria, in which blood is present in urine of the subject, is a presenting clinical symptom of UC. The American Urological Association (AUA) recommends risk stratification for patients with hematuria, and many of these patients may be recommended to undergo a diagnostic cystoscopy to confirm or rule of a UC diagnosis, although confirmed UC rates in patients with hematuria may be low. A high proportion of patients with hematuria who undergo further evaluation, including cystoscopy, may have inconclusive findings, such as atypical cytology or equivocal cystoscopy, requiring further invasive tests (e.g., biopsy). Thus, there exists a need for a test that can accurately risk stratify patients with hematuria as more or less likely to have UC, thereby reducing unnecessary cystoscopies, imaging, and costs, and allowing for timely detection and management of patients with UC. For example, such tests can encompass improved algorithms for determining a presence, absence, or risk of having bladder cancer. Such algorithms may have improved performance metrics, such as sensitivity, specificity, positive predictive value, negative predictive value, and accuracy, as compared to standard-of-care triage and detection.I. METHODS

[0018] Provided herein, in some embodiments, are methods for analyzing genetic data and transcriptomic data derived from a sample of a subject with a trained machine learning algorithm to determine a degree of cancer risk in the subject.

[0019] Methods disclosed herein, in some embodiments, comprise detecting a presence of cancer (e.g., bladder cancer) in the subject, based at least in part on the degree of cancer risk determined by the trained machine learning algorithm.

[0020] Methods disclosed herein, in some embodiments, may comprise stratifying the subject with cancer or suspected of having the cancer (e.g., bladder cancer) as having or not having aAttomey Docket No. 70528-701.601subtype of the cancer (e.g., urothelial bladder cancer), at least in part on the degree of cancer risk determined by the trained machine learning algorithm. The bladder cancer disclosed herein may be confirmed by imaging, a cystoscopy, or some combination thereof. In some instances, the imaging comprises an ultrasound, a renal ultrasound, or an axial upper tract imaging, or any combination thereof. In some instances, the cystoscopy is considered a standard of care in diagnosing UC. In some instances, the cystoscopy is a white light cystoscopy, a flexible cystoscopy, a rigid cystoscopy, or any combination thereof. In some embodiments, hematuria is diagnosed based at least in part on results from a physical examination, a clinical evaluation, a urinalysis, or any combination thereof. A pathological confirmation may also be used to confirm the bladder cancer. In some instances, the pathological confirmation comprises a biopsy of the bladder of the subject. A risk stratification of the subject may be calculated in accordance with American Urological Association (AU A) guidelines. The AUA guidelines, in some cases, determines that the subject is indicated to undergo a cystoscopy to confirm a diagnosis of UC. In some such cases, a diagnosis of UC is ruled out by a cystoscopy.

[0021] Referring to FIG. 4, the methods of the present disclosure involve, in some embodiments, performing or having performed a genotyping assay to determine if the subject has one or more genotypes at one or more polymorphisms; performing or having performed a transcriptomic assay to determine a quantitative measure of one or more transcriptomic markers; analyzing the genotyping and transcriptomic assays; if the subject has a positive cancer result, then administering a cystoscopy exam, and if the patient has a negative cancer result, then administering a cystoscopy exam can be avoided. In some embodiments, the methods of the present disclosure further involve obtaining or having obtained a sample from the subject.

[0022] The subject disclosed herein can be a mammal, such as for example a mouse, rat, guinea pig, rabbit, non-human primate, or farm animal. In some instances, the subject is human. In some instances, the subject is suffering from a symptom related to a disease or condition disclosed herein (e.g., hematuria, pelvic pain, swelling in the legs or feet, urinary incontinence, weight loss, fatigue, loss of appetite, abdominal pain, nausea, and vomiting).

[0023] In some embodiments, the subject has hematuria, defined as blood in the urine of the subject. The hematuria may be gross hematuria (GH), where blood is visibly present in the urine of the subject. The hematuria may be microhematuria (MH), where blood is present in small amount in the urine such that the blood is not visible to a naked eye (e.g., without a microscope). The sample for the subject may be a biological sample. In some embodiments, the sample is a urine sample.

[0024] The genotypes described herein may be detected using a genotyping assay. In some embodiments, the genotyping assay detects the genotype at one or more polymorphisms. InAttomey Docket No. 70528-701.601some embodiments, the genotyping assay detects the genotype at one or more polymorphisms generates a first data set. In some embodiments, the genotypes are heterozygous for a risk allele at the one or more polymorphisms. In some instances, the genotyping assay comprises a polymerase chain reaction (PCR). In some embodiments, the PCR is a quantitative reversetranscription PCR (rt-PCR). In some embodiments, the one or more polymorphisms comprise rsl21913482, rsl21913483, rsl21913479, rsl21913485, rsl242535815, rsl561215364, or a polymorphism in linkage disequilibrium therewith as determined by a coefficient of determination R2of at least 0.85, or any combination thereof.

[0025] The transcriptomic markers described herein may be detected using a transcriptomic assay. In some instances, the transcriptomic assay determines a quantitative measure of the transcriptomic markers. In some embodiments, the quantitative measure of the transcriptomic markers generates a second data set. In some instances, the transcriptomic assay comprises a polymerase chain reaction (PCR). In some instances, the transcriptomic assay comprises a droplet-digital PCR (ddPCR). In some cases, the transcriptomic markers comprise Midkine (MDK), Cyclin Dependent Kinase 1 (CDK1), Insulin Like Growth Factor Binding Protein 5 (IGFBP5), Homeobox A13 (HOXA13), or C-X-C Motif Chemokine Receptor 2 (CXCR2), or any combination thereof.

[0026] The first data set and the second data set described herein may be analyzed using a trained machine learning algorithm. The trained machine learning algorithm may determine an output indicative of a degree of cancer risk of the subject. In some instances, the degree of cancer risk is a categorical cancer risk selected from among a plurality of distinct categorical cancer risks. In some instances, the distinct categorical cancer risks comprises a high cancer risk, an intermediate cancer risk, or a low cancer risk, or any combination thereof. The trained machine learning algorithm may determine an assigned classification. In some instances, the assigned classification is a positive cancer risk. The positive cancer risk classification may be output when the distinct categorical cancer risk of a subject comprises a high cancer risk or the intermediate cancer risk. In some instances, the assigned classification is a negative cancer risk. The negative cancer risk classification may be output when the distinct categorical cancer risk of a subject comprises a low cancer risk.

[0027] The subject may have a positive cancer classification as disclosed herein. In some instances, a cystoscopy is performed to confirm a UC diagnosis of the positive cancer classification. In some embodiments, the method further comprises administering a treatment to the subject, wherein the treatment is capable of treating UC.

[0028] The subject may have a negative cancer classification as disclosed herein. In some instances, the negative cancer classification determines that the subject need not undergo aAttomey Docket No. 70528-701.601cystoscopy for confirmation of UC, even if under other risk stratifiers (e.g., AUA risk stratification) recommend a cystoscopy. Thus, in some cases, where the subject has a negative cancer classification, an unnecessary cystoscopy may be avoided.

[0029] In some embodiments, the method comprises: determining whether the subject has a positive risk of bladder cancer by: obtaining or having obtained a sample from the subject; performing or having performed a genotyping assay to determine if the subject has one or more genotypes at one or more polymorphisms; performing or having performed a transcriptomic assay to determine if a quantitative measure of one or more transcriptomic markers; analyzing results of the genotyping and transcriptomic assays; if the subject has a positive cancer result, then administering a cystoscopy exam, and if the patient has a negative cancer result, then administering a cystoscopy exam can be avoided. In some embodiments, the

[0030] Methods disclosed herein, in some embodiments, comprise monitoring efficacy of a treatment for the cancer (e.g., bladder cancer) in the subject based, at least in part on the degree of cancer risk determined by the trained machine learning algorithm. In some embodiments, the methods comprise administering a therapeutic agent for treating the cancer to the subject; and monitoring the efficacy of the therapeutic agent in treating the cancer by analyzing the genetic data and transcriptomic data from a sample obtained from the subject with the trained machine learning algorithm of the present disclosure.

[0031] Methods disclosed herein, in some embodiments, comprise preparing a sample obtained from a subject suspected of having cancer for determining whether the subject has or does not have the cancer or a subtype of the cancer. The methods comprise: (a) extracting a plurality of nucleic acids from a sample disclosed herein that has been obtained from a subject; and (b) enriching a target nucleic acid from the plurality of nucleic acids comprising one or more polymorphisms disclosed herein, wherein the enriching is performed by (i) brining a fluid reaction formulation comprising a synthetic oligonucleotide molecule in contact with the sample; (ii) hybridizing the synthetic oligonucleotide molecule and the target nucleic acid molecule; and (iii) amplifying the target nucleic acid molecule, thereby enriching the target nucleic acid molecule in the fluid reaction formulation. In some embodiments, the plurality of nucleic acids comprise one or more polymorphisms. In some embodiments, the plurality of nucleic acids comprise one or more transcriptomic markers. In some embodiments, the methods comprise a first plurality of nucleic acids and a second plurality of nucleic acids. In some embodiments, the first plurality of nucleic acids comprise one or more polymorphisms. In some embodiments, second plurality of nucleic acids comprises one or more transcriptomic markers. In some embodiments, the one or more polymorphisms comprise rsl21913482, rsl21913483, rsl21913479, rsl21913485, rsl242535815, rsl561215364, or a polymorphism in linkageAttomey Docket No. 70528-701.601disequilibrium therewith as determined by a coefficient of determination R2of at least 0.85, or any combination thereof. In some embodiments, the one or more transcriptomic markers comprise Midkine (MDK), Cyclin Dependent Kinase 1 (CDK1), Insulin Like Growth Factor Binding Protein 5 (IGFBP5), Homeobox Al 3 (HOXA13), or C-X-C Motif Chemokine Receptor 2 (CXCR2), or any combination thereof.

[0032] To practice the methods and systems provided herein, genetic material may be extracted from a sample obtained from a subject, e.g., a sample of blood or serum. In certain embodiments where nucleic acids are extracted, the nucleic acids are extracted using any technique that does not interfere with subsequent analysis. In certain embodiments, this technique uses alcohol precipitation using ethanol, methanol or isopropyl alcohol. In certain embodiments, this technique uses phenol, chloroform, or any combination thereof. In certain embodiments, this technique uses cesium chloride. In certain embodiments, this technique uses sodium, potassium or ammonium acetate or any other salt commonly used to precipitate DNA.

[0033] Methods disclosed herein, in some embodiments, comprise treating a disease or a condition in a subject suspected of having cancer or a subtype of the cancer based, at least in part, on the degree of cancer risk determined by the trained machine learning algorithm.

[0034] For example, the treatment if the subtype of cancer is identified may be a treatment for urothelial cancer. If the subtype of the cancer is not identified, then further screening may be recommended.

[0035] In some embodiments, the cancer is bladder cancer. In some embodiments, the bladder cancer is a non-muscle invasive bladder cancer, wherein the bladder cancer has not grown into a muscle layer of the bladder. In some embodiments, the bladder cancer is a muscle invasive bladder cancer, wherein the bladder cancer has grown into the muscle layer of the bladder. In some embodiments, the cancer is urothelial carcinoma (UC). UC may also be known as transitional cell carcinoma (TCC). In some embodiments, the bladder cancer may be UC with a divergent differentiation. In some embodiments, the bladder cancer is a squamous cell carcinoma. In some embodiments, the bladder cancer is an adenocarcinoma. In some embodiments, the bladder cancer is a small cell carcinoma. In some embodiments, the bladder cancer is a sarcoma. In some embodiments, the bladder cancer is a non-invasive flat carcinoma (i.e., a carcinoma in situ). In some embodiments, the bladder cancer is a non-invasive papillary carcinoma. In some embodiments, the non-invasive papillary carcinoma is a papillary urothelial neoplasm of low-malignant potential (PUNLMP). In some embodiments, the non-invasive papillary carcinoma is a non-invasive low-grade papillary urothelial carcinoma (LGPUC). In some embodiments, the non-invasive papillary carcinoma is a non-invasive high-grade papillary urothelial carcinoma (HGPUC).Attomey Docket No. 70528-701.601

[0036] Non-limiting examples of treatments for the cancer include a surgical resection, a chemotherapy, an immunotherapy, a radiation therapy, a targeted therapy, or any combination thereof. In some embodiments, the surgical resection is a transurethral resection of the bladder tumor. Non-limiting examples of the chemotherapy include intravesical chemotherapy, gemcitabine, cisplatin, carboplatin, paclitaxel, docetaxel, ifosfamide, doxorubicin, methotrexate, vinblastine, mitomycin, 5 -fluorouracil (5-FU), or any combination thereof. In some embodiments, antibody-drug conjugates may be used.

[0037] In some embodiments, the subject is a human subject. In some instances, the subject is suffering from a symptom related to a disease or condition disclosed herein (e.g., hematuria, pelvic pain, swelling in the legs or feet, urinary incontinence, weight loss, fatigue, loss of appetite, abdominal pain, nausea, and vomiting). In some embodiments, the subject has hematuria, defined as blood in the urine of the subject. The hematuria may be gross hematuria (GH), where blood is visibly present in the urine of the subject. The hematuria may be microhematuria (MH), where blood is present in small amount in the urine such that the blood is not visible to a naked eye (i.e., without a microscope). The sample for the subject may be a biological sample. In some embodiments, the sample is a urine sample.

[0038] The genetic data may include genotypes at one or more polymorphic loci. In some embodiments, the one or more polymorphic loci comprises rsl21913482, rsl21913483, rsl21913479, rsl21913485, rsl242535815, rsl561215364, or a polymorphism in linkage disequilibrium therewith as determined by a coefficient of determination R2of at least 0.85, or any combination thereof. Linkage disequilibrium refers to the non-random association of alleles or indels in different gene loci in a given population. LD may be measured by a D’ value corresponding to the difference between an observed and expected allele or indel frequencies in the population (D=Pab-PaPb), which is scaled by a theoretical maximum value of D. LD may be defined by an R2value corresponding to the difference between an observed and expected unit of risk frequencies in the population (D=Pab-PaPb), which is scaled by the individual frequencies of the different loci. In some embodiments, the linkage disequilibrium may be defined by an R2value of at least 0.80, 0.85, 0.90, 0.95, or 1.0. In some embodiments, the one or more polymorphic loci comprises two or more of rsl21913482, rsl21913483, rsl21913479, rsl21913485, rsl242535815, rsl561215364, or a polymorphism in linkage disequilibrium therewith as determined by a coefficient of determination R2of at least 0.85, or any combination thereof. In some embodiments, the one or more polymorphic loci comprises three or more of rsl21913482, rsl21913483, rsl21913479, rsl21913485, rsl242535815, rsl561215364, or a polymorphism in linkage disequilibrium therewith as determined by a coefficient of determination R2of at least 0.85, or any combination thereof. In some embodiments, the one orAttomey Docket No. 70528-701.601more polymorphic loci comprises four or more of rsl21913482, rsl21913483, rsl21913479, rsl21913485, rsl242535815, rsl561215364, or a polymorphism in linkage disequilibrium therewith as determined by a coefficient of determination R2of at least 0.85, or any combination thereof. In some embodiments, the one or more polymorphic loci comprises five or more of rsl21913482, rsl21913483, rsl21913479, rsl21913485, rsl242535815, rsl561215364, or a polymorphism in linkage disequilibrium therewith as determined by a coefficient of determination R2of at least 0.85, or any combination thereof. In some embodiments, the one or more polymorphic loci comprises six or more of rsl21913482, rsl21913483, rsl21913479, rsl21913485, rsl242535815, rsl561215364, or a polymorphism in linkage disequilibrium therewith as determined by a coefficient of determination R2of at least 0.85, or any combination thereof. In some embodiments, the one or more polymorphisms comprise rsl21913482, rsl21913483, rsl21913479, rsl21913485, rsl242535815, and rsl561215364.

[0039] In some embodiments, the one or more polymorphisms comprise rsl21913482. In some embodiments, rsl21913482 is referred to as MT742 rsl21913482. rsl21913482 is a single nucleotide variant of the fibroblast growth factor receptor 3 (FGFR3) gene. rsl21913482 variations are associated with UC. In some embodiments, the position of rsl21913482 is at chr4: 1801837 (GRCh38.pl4). In some embodiments, the position of rsl21913482 is at chr4: 1803564 (GRCh37). Effects of rsl21913482 variations may include a protein change of R248C.

[0040] In some embodiments, the one or more polymorphisms comprise rsl21913483. In some embodiments, rsl21913483 is referred to as MT746 rsl21913483. rsl21913483 is a single nucleotide variant of the fibroblast growth factor receptor 3 (FGFR3) gene. rsl21913483 variations are associated with UC. In some embodiments, the position of rsl21913483 is at chr4: 1801841 (GRCh38.pl4). In some embodiments, the position of rsl21913483 is at chr4: 1803568 (GRCh37). Effects of rsl21913482 variations may include a protein change of S249F or S249C.

[0041] In some embodiments, the one or more polymorphisms comprise rsl21913479. In some embodiments, rsl21913479 is referred to as MT1114 rsl21913479. rsl21913479 is a single nucleotide variant of the fibroblast growth factor receptor 3 (FGFR3) gene. rsl21913479 variations are associated with UC. In some embodiments, the position of rsl21913479 is at chr4: 1804362 (GRCh38.pl4). In some embodiments, the position of rsl21913479 is at chr4: 1806089 (GRCh37). Effects of rsl21913479 variations may include a protein change of G370S orG370C.

[0042] In some embodiments, the one or more polymorphisms comprise rsl21913485. In some embodiments, rs!21913485 is referred to as MT1124 rs!21913485. rs!21913485 is a singleAttomey Docket No. 70528-701.601nucleotide variant of the fibroblast growth factor receptor 3 (FGFR3) gene. rsl21913485 variations are associated with UC. In some embodiments, the position of rsl21913485 is at chr4: 1804372 (GRCh38.pl4). In some embodiments, the position of rsl21913485 is at chr4: 1806099 (GRCh37). Effects of rsl21913485 variations may include a protein change of Y373C or Y375C.

[0043] In some embodiments, the one or more polymorphisms comprise rsl242535815. In some embodiments, rsl242535815 is referred to as MT124 rsl242535815. rsl242535815 is a single nucleotide variant of the telomerase reverse transcriptase (TERT) gene located in the promotor region. rsl242535815 variations are associated with UC. In some embodiments, the position of rsl242535815 is at chr5:1295113 (GRCh38.pl4). In some embodiments, the position of rsl242535815 is at chr5: 1295228 (GRCh37). Effects of rsl242535815 c.-124C>T mutations may include increases in expression of TERT.

[0044] In some embodiments, the one or more polymorphisms comprise rsl561215364. In some embodiments, rsl561215364 is referred to as MT146 rsl561215364. rsl561215364 is a single nucleotide variant of the telomerase reverse transcriptase (TERT) gene located in the promotor region. rsl561215364 variations are associated with UC. In some embodiments, the position of rsl561215364 is at chr5:1295135 (GRCh38.pl4). In some embodiments, the position of rsl561215364 is at chr5: 1295250 (GRCh37). Effects of rsl561215364 c.-146C>T mutations may induce over expression of proteins.

[0045] In some embodiments, the polymorphic loci include one or polymorphic loci in linkage disequilibrium therewith. Linkage disequilibrium may be defined by an R2value of at least 0.80, 0.85, 0.90, 0.95, or 1.0.

[0046] In some embodiments, polymorphic locus or a genotype at a polymorphic locus is associated with the cancer (e.g., bladder cancer) with a P value of at most about 1.0 x 10'6, about 1.0 x 10'7, about 1.0 x 10'8, about 1.0 x 10'9, about 1.0 x 10'10, about 1.0 x IO'20, about 1.0 x 10"30, about 1.0 x IO'40, about 1.0 x 10'50, about 1.0 x IO'60, about 1.0 x IO'70, about 1.0 x IO'80, about 1.0 x IO'90, or about 1.0 x IO'100. In some embodiments, polymorphic locus or a genotype at a polymorphic locus is protective for the cancer (e.g., bladder cancer) with a P value of at most about 1.0 x 10'6, about 1.0 x 10'7, about 1.0 x 10'8, about 1.0 x 10'9, about 1.0 x 10'10, about 1.0 x IO'20, about 1.0 x 10'30, about 1.0 x IO'40, about 1.0 x 10'50, about 1.0 x IO'60, about 1.0 x IO'70, about 1.0 x IO'80, about 1.0 x IO'90, or about 1.0 x IO'100.

[0047] The transcriptomic data may include an amount or a level of one or more transcriptomic markers. In some embodiments, the one or more transcriptomic markers comprise Midkine (MDK), Cyclin Dependent Kinase 1 (CDK1), Insulin Like Growth Factor Binding Protein 5 (IGFBP5), Homeobox A13 (HOXA13), or C-X-C Motif Chemokine Receptor 2 (CXCR2), orAttomey Docket No. 70528-701.601any combination thereof. In some embodiments, the one or more transcriptomic markers comprises two or more of Midkine (MDK), Cyclin Dependent Kinase 1 (CDK1), Insulin Like Growth Factor Binding Protein 5 (IGFBP5), Homeobox A13 (HOXA13), or C-X-C Motif Chemokine Receptor 2 (CXCR2), or any combination thereof. In some embodiments, the one or more transcriptomic markers comprises three or more of Midkine (MDK), Cyclin Dependent Kinase 1 (CDK1), Insulin Like Growth Factor Binding Protein 5 (IGFBP5), Homeobox Al 3 (HOXA13), or C-X-C Motif Chemokine Receptor 2 (CXCR2), or any combination thereof. In some embodiments, the one or more transcriptomic markers comprises four or more of Midkine (MDK), Cyclin Dependent Kinase 1 (CDK1), Insulin Like Growth Factor Binding Protein 5 (IGFBP5), Homeobox A13 (HOXA13), or C-X-C Motif Chemokine Receptor 2 (CXCR2), or any combination thereof. In some embodiments, the one or more transcriptomic markers comprise Midkine (MDK), Cyclin Dependent Kinase 1 (CDK1), Insulin Like Growth Factor Binding Protein 5 (IGFBP5), Homeobox Al 3 (HOXA13), and C-X-C Motif Chemokine Receptor 2 (CXCR2).

[0048] In some embodiments, the one or more transcriptomic markers comprise Midkine (MDK). MDK is a multifunctional protein that forms a family of heparin-binding growth factors. MDK may be overexpressed in many types of cancer, including bladder cancer and UC. Overexpression of MDK may be especially noticeable during progression of a tumor into more advanced stages. MDK expression in tumors may be determined by multiple forms of analysis, including blood and urine analysis.

[0049] In some embodiments, the one or more transcriptomic markers comprise Cyclin Dependent Kinase 1 (CDK1). CDK1 is kinase that is crucial in cell cycle progression. One indicators of malignancy, such as unrestricted cell proliferation, may be caused by alterations in CDK1 activity.

[0050] In some embodiments, the one or more transcriptomic markers comprise Insulin Like Growth Factor Binding Protein 5 (IGFBP5). IGFBP5 is associated with several cancers, including bladder cancer and UC.

[0051] In some embodiments, the one or more transcriptomic markers comprise Homeobox Al 3 (HOXA13). HOXA13 is associated with several cancers, including bladder cancer and UC.

[0052] In some embodiments, the one or more transcriptomic markers comprise C-X-C Motif Chemokine Receptor 2 (CXCR2). CXCR2 is an inflammation marker.

[0053] In some embodiments, the amount of the transcriptomic marker is an absolute amount, whereas a level of the transcriptomic marker may be an expression level that is relative to a reference expression level for the transcriptomic marker. In some embodiments, a composite score comprises the amount of the transcriptomic marker. In some embodiments, the compositeAttomey Docket No. 70528-701.601score is compared to a threshold to determine a level of risk. In some embodiments, the level of risk is high. In some embodiments, the level of risk is intermediate. In some embodiments, the level of risk is low. In some embodiments, the threshold to determine a high risk is above 2.5. In some embodiments, the threshold to determine an intermediate risk is between 1.46 and 2.5. In some embodiments, the threshold to determine a high risk is below 1.46.A. Methods of Detection

[0054] Disclosed herein, in some embodiments, are methods of assaying the sample with a genotyping assay. An assaying device that may be useful for the methods herein include a quantitative polymerase chain reaction (qPCR), nucleic acid sequencing reaction, or a genotyping array, polymerase chain reactions (PCRs), or droplet-digital PCRs (ddPCR). In some embodiments, the ddPCR utilizes a Taq polymerase in a standard PCR reaction to amplify a target DNA fragment from a sample. ddPCR may also partition the PCR reaction into thousands of individual reaction vessels before an amplification step. ddPCR may also comprise acquiring data at a reaction end point.

[0055] Disclosed herein, in some embodiments, are methods of assaying the sample with a transcriptomic assay. An assaying device that may be useful for the methods herein include a polymerase chain reaction (PCR), a quantitative polymerase chain reaction (qPCR), a quantitative reverse-transcription PCR (rtPCR), nucleic acid sequencing reaction, a genotyping array, or droplet-digital PCRs (ddPCR). In some embodiments, the assay device comprises a quantitative rtPCR. In some embodiments, the quantitative rtPCR is performed in one step, wherein a reverse transcription step and a PCR step are performed in a single tube and buffer, using a reverse transcriptase along with a DNA polymerase, wherein only sequence-specific primers are utilized. In some embodiments, the quantitative rtPCR is performed in two steps, wherein the reverse transcription step and the PCR step are performed in separate tubes comprising different optimized buffers, reaction conditions, and priming strategies.

[0056] Disclosed herein, in some embodiments, are methods of sample collection. In some embodiments, the method comprises a step of collecting a sample. In some embodiments, collecting a sample comprises collecting a urine sample from a subject with hematuria. In some embodiments, collecting a sample comprises a midstream urine sample. In some embodiments, collecting a sample comprises a midstream urine sample prior to a cystoscopy. In some embodiments, collecting a sample comprises a midstream urine sample at least 14 days after a subject had a cystoscopy.

[0057] Disclosed herein, in some embodiments, are methods of sample preparation. Methods of sample preparation disclosed herein, in some embodiments, comprise preparing a sampleAttomey Docket No. 70528-701.601obtained from a subject suspected of having cancer for determining whether the subject has or does not have the cancer or a subtype of the cancer. The methods comprise: (a) extracting a plurality of nucleic acids from a sample disclosed herein that has been obtained from a subject; and (b) enriching a target nucleic acid from the plurality of nucleic acids comprising one or more polymorphisms disclosed herein, wherein the enriching is performed by (i) brining a fluid reaction formulation comprising a synthetic oligonucleotide molecule in contact with the sample; (ii) hybridizing the synthetic oligonucleotide molecule and the target nucleic acid molecule; and (iii) amplifying the target nucleic acid molecule, thereby enriching the target nucleic acid molecule in the fluid reaction formulation. In some embodiments, the plurality of nucleic acids comprise one or more polymorphisms. In some embodiments, the plurality of nucleic acids comprise one or more transcriptomic markers. In some embodiments, the methods comprise a first plurality of nucleic acids and a second plurality of nucleic acids. In some embodiments, the first plurality of nucleic acids comprise one or more polymorphisms. In some embodiments, second plurality of nucleic acids comprises one or more transcriptomic markers. In some embodiments, the one or more polymorphisms comprise rsl21913482, rsl21913483, rsl21913479, rsl21913485, rsl242535815, rsl561215364, or a polymorphism in linkage disequilibrium therewith as determined by a coefficient of determination R2 of at least 0.85, or any combination thereof. In some embodiments, the one or more transcriptomic markers comprise Midkine (MDK), Cyclin Dependent Kinase 1 (CDK1), Insulin Like Growth Factor Binding Protein 5 (IGFBP5), Homeobox Al 3 (HOXA13), or C-X-C Motif Chemokine Receptor 2 (CXCR2), or any combination thereof.

[0058] In some embodiments, the sample is collected from a subject that has hematuria, defined as blood in the urine of the subject. The hematuria may be gross hematuria (GH), where blood is visibly present in the urine of the subject. The hematuria may be microhematuria (MH), where blood is present in small amount in the urine such that the blood is not visible to a naked eye (i.e., without a microscope). The sample for the subject may be a biological sample. In some embodiments, the sample is a urine sample.B. Computer-Implemented Methods

[0059] Disclosed herein are computer-implemented methods for (i) training and testing a trained algorithm, (ii) using the trained algorithm to process data to determine a presence, absence, or risk of bladder cancer of a subject, (iii) determining a quantitative measure indicative of a presence, absence, or risk of bladder cancer of a subject, (iv) identifying or monitoring the bladder cancer of the subject, and / or (v) electronically outputting a report that indicative of a presence, absence, or risk of bladder cancer of the subject.Attomey Docket No. 70528-701.601

[0060] A trained algorithm may be used to process datasets to determine an output indicative of a presence, absence, or risk of bladder cancer. For example, the datasets may be generated by assaying biological samples derived from the subject, such as to determine quantitative measures of sequences at each of a plurality of bladder cancer-associated genomic loci in the biological samples. The trained algorithm may be trained to identify the bladder cancer with a certain accuracy, such as an accuracy of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or more than 99%.

[0061] The trained algorithm may comprise a supervised machine learning algorithm. The trained algorithm may comprise a classification and regression tree (CART) algorithm. The supervised machine learning algorithm may comprise, for example, a Random Forest, a support vector machine (SVM), a neural network, or a deep learning algorithm. The trained algorithm may comprise a differential expression algorithm. The differential expression algorithm may comprise a use comparison of stochastic models, generalized Poisson (GPseq), mixed Poisson (TSPM), Poisson log-linear (PoissonSeq), negative binomial (edgeR, DESeq, baySeq, NBPSeq), linear model fit by MAANOVA, or a combination thereof. The trained algorithm may comprise an unsupervised machine learning algorithm.

[0062] The methods and systems described herein may comprise one or more machine learning techniques. In some cases, machine learning may comprise identifying and recognizing patterns in existing data in order to facilitate making predictions for subsequent data. Machine learning may include a machine learning model (which may include, for example, a machine learning algorithm). Machine learning, whether analytical or statistical in nature, may provide deductive or abductive inference based on real or simulated data. The machine learning model may be a trained machine learning model. Machine learning techniques may comprise one or more supervised, semi-supervised, self-supervised, or unsupervised machine learning techniques. For example, a machine learning model may comprise a trained machine learning model that is trained through supervised learning (e.g., various parameters are determined as weights, scaling factors, or polynomial coefficients).

[0063] Machine learning may comprise one or more of regression analysis, regularization, classification, dimensionality reduction, ensemble learning, meta learning, association rule learning, cluster analysis, anomaly detection, deep learning, or ultra-deep learning, machine learning may comprise: k-means, k-means clustering, k-nearest neighbors, learning vector quantization, linear regression, non-linear regression, least squares regression, partial least squares regression, logistic regression, stepwise regression, multivariate adaptive regressionAttomey Docket No. 70528-701.601splines, ridge regression, principal component regression, least absolute shrinkage and selection operation (LASSO), least angle regression, canonical correlation analysis, factor analysis, independent component analysis, linear discriminant analysis, multidimensional scaling, nonnegative matrix factorization, principal components analysis, principal coordinates analysis, projection pursuit, Sammon mapping, t-distributed stochastic neighbor embedding, AdaBoosting, boosting, gradient boosting, bootstrap aggregation, ensemble averaging, decision trees, conditional decision trees, boosted decision trees, gradient boosted decision trees, random forests, stacked generalization, Bayesian networks, Bayesian belief networks, naive Bayes, Gaussian naive Bayes, multinomial naive Bayes, hidden Markov models, hierarchical hidden Markov models, support vector machines, encoders, decoders, auto-encoders, stacked autoencoders, perceptrons, multi-layer perceptrons, artificial neural networks, feedforward neural networks, convolutional neural networks, recurrent neural networks, residual neural networks, physics-informed neural networks, long short-term memory, deep belief networks, deep Boltzmann machines, deep convolutional neural networks, deep recurrent neural networks, large language models, transformer models, vision transformers, or generative adversarial networks.

[0064] In some embodiments, the trained machine learning algorithm comprises a boosted regression tree (BRT), a Bayesian Additive Regression Tree (BART), a random forest (RF) algorithm, a boosted tree algorithm, a gradient boosting machine (GBM), a Bayesian Model Averaging (BMA) with a decision tree, a classification and regression tree (CART) model, a support vector machine (SVM), a Linear Discriminate Analysis (LDA), a Logistic Regression (LogReg), a K-nearest 5 neighbors (Kn5n), a partition tree classifier (TREE), or a partially fixed BART.

[0065] In some embodiments, the trained machine learning algorithm is a single trained machine learning algorithm. In some embodiments, the trained machine learning algorithm is an integrated algorithm or an ensemble machine learning algorithm comprising two or more algorithms. For example, a set of outputs from different algorithms may be processed together to generate an overall output for the trained machine learning algorithm. As another example, different algorithms may be concatenated together, such that the output of a first algorithm maybe fed into the input of a second algorithm.

[0066] In some embodiments, the trained machine learning algorithm comprises an inverse logit function. For example, the inverse logit function may be an inverse logit of a second-order polynomial. The second-order polynomial comprises coefficients obtained by fitting a model such as a logistic regression model. The inverse logit function may determine an estimated probability of the subject having cancer. One or more thresholds may be applied to the estimatedAttomey Docket No. 70528-701.601probability to generate a categorical measure of cancer risk (e.g., high risk, intermediate risk, or low risk.

[0067] In some cases, training the machine learning model may include selecting one or more untrained data models to train using a training dataset set. The selected untrained data models may include any type of untrained machine learning models for supervised, semi-supervised, self-supervised, or unsupervised machine learning. The selected untrained data models may be specified based upon input (e.g., user input) specifying relevant parameters to use as predicted variables or other variables to use as potential explanatory variables. For example, the selected untrained data models may be specified to generate an output (e.g., a prediction) based upon the input. Conditions for training the machine learning model from the selected untrained data models may likewise be selected, such as limits on the machine learning model complexity or limits on the machine learning model refinement past a certain point. The machine learning model may be trained (e.g., via a computer system such as a server) using the training dataset set. In some cases, a first subset of the training dataset set may be selected to train the machine learning model. The selected untrained data models may then be trained on the first subset of training dataset set using appropriate machine learning techniques, based upon the type of machine learning model selected and any conditions specified for training the machine learning model. In some cases, due to the processing power requirements of training the machine learning model, the selected untrained data models may be trained using additional computing resources (e.g., cloud computing resources). Such training may continue, in some cases, until at least one aspect of the machine learning model is validated and meets selection criteria to be used as a predictive model.

[0068] In some cases, the machine learning model may be validated using a second subset of the training dataset set (e.g., distinct from the first subset of the training dataset set) to determine accuracy and robustness of the machine learning model. Such validation may include applying the machine learning model to the second subset of the training dataset set to make predictions derived from the second subset of the training dataset. The machine learning model may then be evaluated to determine whether performance is sufficient based upon the derived predictions. The sufficiency criteria applied to the machine learning model may vary depending upon the size of the training dataset set available for training, the performance of previous iterations of trained models, or user-specified performance requirements. If the machine learning model does not achieve sufficient performance, additional training may be performed. Additional training may include refinement of the machine learning model or retraining on a different first subset of the training dataset, after which the new machine learning model may again be validated and assessed. When the machine learning model has achieved sufficient performance, in some cases,Attomey Docket No. 70528-701.601the machine learning may be stored for present or future use. The machine learning model may be stored as sets of parameter values or weights for analysis of further input (e.g., further relevant parameters to use as further predicted variables, further explanatory variables, further user interaction data, etc.), which may also include analysis logic or indications of model validity in some instances. In some cases, a plurality of machine learning models may be stored for generating predictions under different sets of input data conditions. In some embodiments, the machine learning model may be stored in a database (e.g., associated with a server).

[0069] The machine learning model may implement a decision tree. A decision tree may be a supervised machine learning algorithm that can be applied to both regression and classification problems. For example, a decision tree may grow from a root (base condition), and when it meets a condition (internal node / feature), it may split into multiple branches. The end of the branch that does not split anymore may be an outcome (leaf). A decision tree can be generated using a training dataset set according to the following operations: (A) starting from a root node (the entire dataset), the algorithm may split the dataset in two branches using a decision rule or branching criterion; (B) each of these two branches may generate a new child node; (C) for each new child node, the branching process may be repeated until the dataset cannot be split any further; (D) each branching criterion may be chosen to maximize information gain (e.g., a quantification of how much a branching criterion reduces a quantification of how mixed the labels are in the children nodes). The labels may be the data or the classification that is predicted by the decision tree.

[0070] A random forest regression is an extension of the decision tree model that may yield more robust predictions by stretching the use of the training dataset partition. Whereas a decision tree may make a single pass through the data, a random forest regression may bootstrap 50% of the data (e.g., with replacement) and build many trees. Rather than using all explanatory variables as candidates for splitting, a random subset of candidate variables may be used for splitting, which may enable trees that have different data and different variables (hence the term random). The predictions from the trees, which may be collectively referred to as the “forest,” may then be averaged to produce a final prediction. Many trees (e.g., ten trees, fifty trees, one hundred trees, one thousand trees, etc.) may be included in a random forest model, with a number (e.g., 3, 6, 10, etc.) of terms sampled per split, a minimum of number (e.g., 1, 2, 4, 10, etc.) of splits per tree, and a minimum split size (e.g., 16, 32, 64, 128, 256, etc.). Random forests may be trained in a similar way as decision trees. Specifically, training a random forest may include the following operations: (A) randomly select k features from the total number of features; (B) create a decision tree from these k features using the same operations as forAttomey Docket No. 70528-701.601generating a decision tree; and (C) repeat the previous two operations until a target number of trees is created.

[0071] A random forest classifier, which may comprise a plurality of decision trees where the output prediction may be the mode of the predicted classifications of the individual trees, can be helpful in reducing overfitting to training dataset. In some cases, an ensemble of decision trees can be constructed using a random subset of features at each split or decision node. The Gini criterion may be employed, in some cases, to choose the best partition, where decision nodes having the lowest calculated Gini impurity index are selected. The Gini impurity can be used, in some cases, as a criterion to find informative features based on which the splits in each decision tree may be constructed.

[0072] In some cases, each decision tree of a random forest may comprise one or more decision nodes, where each decision node specifies a predicate condition. For example, decision node may predicate the condition that, for a given dataset, the outcome to an question is a specific outcome. At each decision node, a decision tree can be split based on whether the predicate condition attached to the decision node holds true, leading to various prediction nodes. Each prediction node can comprise output values that represent “votes” for one or more of the classifications or conditions being evaluated by the assessment model. At prediction time, a “vote” can be taken over all of the decision trees, and the majority vote (or mode of the predicted classifications) can be output as the predicted classification.

[0073] In some cases, when the dataset being queried in the assessment model reaches a “leaf’, or a final prediction node with no further downstream splits, the output values of the leaf can be output as the votes for the particular decision tree. Since a random forest model comprises a plurality of decision trees, the final votes across all trees in the forest can be summed to yield the final votes and the corresponding classification of the subject. A large number of decision trees can help reduce overfitting of the assessment model to the training dataset, by reducing the variance of each individual decision tree. For example, an assessment model can comprise, for example, at least about 3 decision trees, at least about 5 decision trees, at least about 10 decision trees, at least about 20 decision trees, at least about 50 decision trees, at least about 100 decision trees, etc.

[0074] The trained algorithm may be configured to accept a plurality of input variables and to produce one or more output values based on the plurality of input variables. The plurality of input variables may comprise one or more datasets indicative of a bladder cancer. For example, an input variable may comprise a number of sequences corresponding to or aligning to each of the plurality of bladder cancer-associated genomic loci. The plurality of input variables may also include clinical health data of a subject.Attomey Docket No. 70528-701.601

[0075] The trained algorithm may comprise a classifier, such that each of the one or more output values comprises one of a fixed number of possible values (e.g., a linear classifier, a logistic regression classifier, etc.) indicating a classification of the biological sample by the classifier. The trained algorithm may comprise a binary classifier, such that each of the one or more output values comprises one of two values (e.g., {0, 1}, {positive, negative}, or {high-risk, low-risk}) indicating a classification of the biological sample by the classifier. The trained algorithm may be another type of classifier, such that each of the one or more output values comprises one of more than two values (e.g., {0, 1, 2}, {positive, negative, or indeterminate}, or {high-risk, intermediate-risk, or low-risk}) indicating a classification of the biological sample by the classifier. The output values may comprise descriptive labels, numerical values, or a combination thereof. Some of the output values may comprise descriptive labels. Such descriptive labels may provide an identification or indication of the disease or disorder state of the subject, and may comprise, for example, positive, negative, high-risk, intermediate-risk, low-risk, or indeterminate. Such descriptive labels may provide an identification of a treatment for the subject’s bladder cancer, and may comprise, for example, a therapeutic intervention, a duration of the therapeutic intervention, and / or a dosage of the therapeutic intervention suitable to treat a bladder cancer. Such descriptive labels may provide an identification of secondary clinical tests that may be appropriate to perform on the subject, and may comprise, for example, an imaging test, a blood test, a computed tomography (CT) scan, a magnetic resonance imaging (MRI) scan, an ultrasound scan, a chest X-ray, a positron emission tomography (PET) scan, a PET-CT scan, a biological cytology, or any combination thereof. For example, such descriptive labels may provide a prognosis of the bladder cancer of the subject. As another example, such descriptive labels may provide a relative assessment of the bladder cancer (e.g., an estimated grade or severity) of the subject. Some descriptive labels may be mapped to numerical values, for example, by mapping “positive” to 1 and “negative” to 0.

[0076] Some of the output values may comprise numerical values, such as binary, integer, or continuous values. Such binary output values may comprise, for example, {0, 1}, {positive, negative}, or {high-risk, low-risk}. Such integer output values may comprise, for example, {0, 1, 2}. Such continuous output values may comprise, for example, a probability value of at least 0 and no more than 1. Such continuous output values may comprise, for example, an unnormalized probability value of at least 0. Such continuous output values may indicate a prognosis of the bladder cancer of the subject. Some numerical values may be mapped to descriptive labels, for example, by mapping 1 to “positive” and 0 to “negative.”

[0077] Some of the output values may be assigned based on one or more cutoff values. For example, a binary classification of samples may assign an output value of “positive” or 1 if theAttomey Docket No. 70528-701.601sample indicates that the subject has at least a 50% probability of having a bladder cancer. For example, a binary classification of samples may assign an output value of “negative” or 0 if the sample indicates that the subject has less than a 50% probability of having a bladder cancer. In this case, a single cutoff value of 50% is used to classify samples into one of the two possible binary output values. Examples of single cutoff values may include about 1%, about 2%, about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, and about 99%.

[0078] As another example, a classification of samples may assign an output value of “positive” or 1 if the sample indicates that the subject has a probability of having a bladder cancer of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or more. The classification of samples may assign an output value of “positive” or 1 if the sample indicates that the subject has a probability of having a bladder cancer of more than about 50%, more than about 55%, more than about 60%, more than about 65%, more than about 70%, more than about 75%, more than about 80%, more than about 85%, more than about 90%, more than about 91%, more than about 92%, more than about 93%, more than about 94%, more than about 95%, more than about 96%, more than about 97%, more than about 98%, or more than about 99%.

[0079] The classification of samples may assign an output value of “negative” or 0 if the sample indicates that the subject has a probability of having a bladder cancer of less than about 50%, less than about 45%, less than about 40%, less than about 35%, less than about 30%, less than about 25%, less than about 20%, less than about 15%, less than about 10%, less than about 9%, less than about 8%, less than about 7%, less than about 6%, less than about 5%, less than about 4%, less than about 3%, less than about 2%, or less than about 1%. The classification of samples may assign an output value of “negative” or 0 if the sample indicates that the subject has a probability of having a bladder cancer of no more than about 50%, no more than about 45%, no more than about 40%, no more than about 35%, no more than about 30%, no more than about 25%, no more than about 20%, no more than about 15%, no more than about 10%, no more than about 9%, no more than about 8%, no more than about 7%, no more than about 6%, no more than about 5%, no more than about 4%, no more than about 3%, no more than about 2%, or no more than about 1%.Attorney Docket No. 70528-701.601

[0080] The classification of samples may assign an output value of “indeterminate” or 2 if the sample is not classified as “positive”, “negative”, 1, or 0. In this case, a set of two cutoff values is used to classify samples into one of the three possible output values. Examples of sets of cutoff values may include {1%, 99%}, {2%, 98%}, {5%, 95%}, {10%, 90%}, {15%, 85%}, {20%, 80%}, {25%, 75%}, {30%, 70%}, {35%, 65%}, {40%, 60%}, and {45%, 55%}.Similarly, sets of n cutoff values may be used to classify samples into one of n+1 possible output values, where n is any positive integer.

[0081] The trained algorithm may be trained with a plurality of independent training samples. Each of the independent training samples may comprise a biological sample from a subject, associated datasets obtained by assaying the biological sample (as described elsewhere herein), and one or more known output values corresponding to the biological sample (e.g., a clinical diagnosis, prognosis, absence, or treatment efficacy of a bladder cancer of the subject).Independent training samples may comprise biological samples and associated datasets and outputs obtained or derived from a plurality of different subjects. Independent training samples may comprise biological samples and associated datasets and outputs obtained at a plurality of different time points from the same subject (e.g., on a regular basis such as weekly, biweekly, or monthly). Independent training samples may be associated with presence of the bladder cancer (e.g., training samples comprising biological samples and associated datasets and outputs obtained or derived from a plurality of subjects known to have the bladder cancer). Independent training samples may be associated with absence of the bladder cancer (e.g., training samples comprising biological samples and associated datasets and outputs obtained or derived from a plurality of subjects who are known to not have a previous diagnosis of the bladder cancer or who have received a negative test result for the bladder cancer).

[0082] The trained algorithm may be trained with at least about 5, at least about 10, at least about 15, at least about 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 100, at least about 150, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, or at least about 500 independent training samples. The independent training samples may comprise biological samples associated with presence of the bladder cancer and / or biological samples associated with absence of the bladder cancer. The trained algorithm may be trained with no more than about 500, no more than about 450, no more than about 400, no more than about 350, no more than about 300, no more than about 250, no more than about 200, no more than about 150, no more than about 100, or no more than about 50 independent training samples associated with presence of the bladder cancer. In some embodiments, the biological sample is independent of samples used to train the trained algorithm.Attomey Docket No. 70528-701.601

[0083] The trained algorithm may be trained with a first number of independent training samples associated with presence of the bladder cancer and a second number of independent training samples associated with absence of the bladder cancer. The first number of independent training samples associated with presence of the bladder cancer may be no more than the second number of independent training samples associated with absence of the bladder cancer. The first number of independent training samples associated with presence of the bladder cancer may be equal to the second number of independent training samples associated with absence of the bladder cancer. The first number of independent training samples associated with presence of the bladder cancer may be greater than the second number of independent training samples associated with absence of the bladder cancer.

[0084] The trained algorithm may be configured to identify the bladder cancer at an accuracy of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or more; for at least about 5, at least about 10, at least about 15, at least about 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 100, at least about 150, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, or at least about 500 independent training samples. The accuracy of identifying the bladder cancer by the trained algorithm may be calculated as the percentage of independent test samples (e.g., subjects known to have the bladder cancer or subjects with negative clinical test results for the bladder cancer) that are correctly identified or classified as having or not having the bladder cancer.

[0085] The trained algorithm may be configured to identify the bladder cancer with a positive predictive value (PPV) of at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or more. The PPV of identifying the bladder cancer using the trained algorithm may be calculated as the percentage of biological samples identified orAttomey Docket No. 70528-701.601classified as having the bladder cancer that correspond to subjects that truly have the bladder cancer.

[0086] The trained algorithm may be configured to identify the bladder cancer with a negative predictive value (NPV) of at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or more. The NPV of identifying the bladder cancer using the trained algorithm may be calculated as the percentage of biological samples identified or classified as not having the bladder cancer that correspond to subjects that truly do not have the bladder cancer.

[0087] The trained algorithm may be configured to identify the bladder cancer with a clinical sensitivity at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.1%, at least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, at least about 99.99%, at least about 99.999%, or more. The clinical sensitivity of identifying the bladder cancer using the trained algorithm may be calculated as the percentage of independent test samples associated with presence of the bladder cancer (e.g., subjects known to have the bladder cancer) that are correctly identified or classified as having the bladder cancer.

[0088] The trained algorithm may be configured to identify the bladder cancer with a clinical specificity of at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, atAttomey Docket No. 70528-701.601least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.1%, at least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, at least about 99.99%, at least about 99.999%, or more. The clinical specificity of identifying the bladder cancer using the trained algorithm may be calculated as the percentage of independent test samples associated with absence of the bladder cancer (e.g., subjects with negative clinical test results for the bladder cancer) that are correctly identified or classified as not having the bladder cancer.

[0089] The trained algorithm may be configured to identify the bladder cancer with an Area-Under-Curve (AUC) of at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.81, at least about 0.82, at least about 0.83, at least about 0.84, at least about 0.85, at least about 0.86, at least about 0.87, at least about 0.88, at least about 0.89, at least about 0.90, at least about 0.91, at least about 0.92, at least about 0.93, at least about 0.94, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, at least about 0.99, or more. The AUC may be calculated as an integral of the Receiver Operator Characteristic (ROC) curve (e.g., the area under the ROC curve) associated with the trained algorithm in classifying biological samples as having the bladder cancer or not having the bladder cancer.

[0090] The trained algorithm may be adjusted or tuned to improve one or more of the performance, accuracy, PPV, NPV, clinical sensitivity, clinical specificity, or AUC of identifying the bladder cancer. The trained algorithm may be adjusted or tuned by adjusting parameters of the trained algorithm (e.g., a set of cutoff values used to classify a biological sample as described elsewhere herein, or weights of a neural network or function such as polynomial function). The trained algorithm may be adjusted or tuned continuously during the training process or after the training process has completed.

[0091] After the trained algorithm is initially trained, a subset of the inputs may be identified as most influential or most important to be included for making high-quality classifications. For example, a subset of the plurality of bladder cancer-associated genomic loci may be identified as most influential or most important to be included for making high-quality classifications or identifications of bladder cancers (or subtypes of bladder cancers). The plurality of bladder cancer-associated genomic loci or a subset thereof may be ranked based on classification metrics indicative of each genomic locus’s influence or importance toward making high-quality classifications or identifications of bladder cancers (or subtypes of bladder cancers). Such metrics may be used to reduce, in some cases significantly, the number of input variables (e.g.,Attomey Docket No. 70528-701.601predictor variables) that may be used to train the trained algorithm to a desired performance level (e.g., based on a desired minimum accuracy, PPV, NPV, clinical sensitivity, clinical specificity, AUC, or a combination thereof). For example, if training the trained algorithm with a plurality comprising several dozen or hundreds of input variables in the trained algorithm results in an accuracy of classification of more than 99%, then training the trained algorithm instead with only a selected subset of no more than about 5, no more than about 10, no more than about 15, no more than about 20, no more than about 25, no more than about 30, no more than about 35, no more than about 40, no more than about 45, no more than about 50, or no more than about 100 such most influential or most important input variables among the plurality can yield decreased but still acceptable accuracy of classification (e.g., at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%). The subset may be selected by rank-ordering the entire plurality of input variables and selecting a predetermined number (e.g., no more than about 5, no more than about 10, no more than about 15, no more than about 20, no more than about 25, no more than about 30, no more than about 35, no more than about 40, no more than about 45, no more than about 50, or no more than about 100) of input variables with the best classification metrics.

[0092] After using a trained algorithm to process the dataset, the bladder cancer may be identified or monitored in the subject. The identification may be based at least in part on quantitative measures of sequence reads of the dataset at a panel of bladder cancer-associated genomic loci (e.g., quantitative measures of RNA transcripts or DNA at the bladder cancer-associated genomic loci), proteomic data comprising quantitative measures of proteins of the dataset at a panel of bladder cancer-associated proteins, and / or metabolome data comprising quantitative measures of a panel of bladder cancer-associated metabolites.

[0093] The bladder cancer may be identified in the subject at an accuracy of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or more. The accuracy of identifying the bladder cancer by the trained algorithm may be calculated as the percentage of independent test samples (e.g., subjects knownAttomey Docket No. 70528-701.601to have the bladder cancer or subjects with negative clinical test results for the bladder cancer) that are correctly identified or classified as having or not having the bladder cancer.

[0094] The bladder cancer may be identified in the subject with a positive predictive value (PPV) of at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or more. The PPV of identifying the bladder cancer using the trained algorithm may be calculated as the percentage of biological samples identified or classified as having the bladder cancer that correspond to subjects that truly have the bladder cancer.

[0095] The bladder cancer may be identified in the subject with a negative predictive value (NPV) of at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or more. The NPV of identifying the bladder cancer using the trained algorithm may be calculated as the percentage of biological samples identified or classified as not having the bladder cancer that correspond to subjects that truly do not have the bladder cancer.

[0096] The bladder cancer may be identified in the subject with a clinical sensitivity of at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.1%, at least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least aboutAttomey Docket No. 70528-701.60199.9%, at least about 99.99%, at least about 99.999%, or more. The clinical sensitivity of identifying the bladder cancer using the trained algorithm may be calculated as the percentage of independent test samples associated with presence of the bladder cancer (e.g., subjects known to have the bladder cancer) that are correctly identified or classified as having the bladder cancer.

[0097] The bladder cancer may be identified in the subject with a clinical specificity of at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.1%, at least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, at least about 99.99%, at least about 99.999%, or more. The clinical specificity of identifying the bladder cancer using the trained algorithm may be calculated as the percentage of independent test samples associated with absence of the bladder cancer (e.g., subjects with negative clinical test results for the bladder cancer) that are correctly identified or classified as not having the bladder cancer.

[0098] In some embodiments, the trained algorithm that is trained on samples independent of the biological sample may be used to determine that the subject is at risk of bladder cancer at an accuracy of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or more.

[0099] After the bladder cancer is identified in a subject, a subtype of the bladder cancer (e.g., selected from among a plurality of subtypes of the bladder cancer) may further be identified. The subtype of the bladder cancer may be determined based at least in part on the quantitative measures of sequence reads of the dataset at a panel of bladder cancer-associated genomic loci (e.g., quantitative measures of RNA transcripts or DNA at the bladder cancer-associated genomic loci), proteomic data comprising quantitative measures of proteins of the dataset at a panel of bladder cancer-associated proteins, and / or metabolome data comprising quantitative measures of a panel of bladder cancer-associated metabolites. For example, the subject may beAttomey Docket No. 70528-701.601identified as being at risk of a subtype of bladder cancer (e.g., selected from among a plurality of subtypes of bladder cancer). After identifying the subject as being at risk of a subtype of bladder cancer, a clinical intervention for the subject may be selected based at least in part on the subtype of bladder cancer for which the subject is identified as being at risk. In some embodiments, the clinical intervention is selected from a plurality of clinical interventions (e.g., clinically indicated for different subtypes of bladder cancer).

[0100] In some embodiments, the trained algorithm may determine that the subject is at risk of bladder cancer of at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or more.

[0101] The trained algorithm may determine that the subject is at risk of bladder cancer at an accuracy of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.1%, at least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, at least about 99.99%, at least about 99.999%, or more.

[0102] Upon identifying the subject as having the bladder cancer or having an elevated risk of having bladder cancer, the subject may be optionally provided with a therapeutic intervention (e.g., prescribing an appropriate course of treatment to treat the bladder cancer of the subject or prophylactically reduce elevated risk of bladder cancer). The therapeutic intervention may comprise a prescription of an effective dose of a drug, a further testing or evaluation of the bladder cancer, a further monitoring of the bladder cancer, or a combination thereof. If the subject is currently being treated for the bladder cancer with a course of treatment, the therapeutic intervention may comprise a subsequent different course of treatment (e.g., to increase treatment efficacy due to non-efficacy of the current course of treatment).Attomey Docket No. 70528-701.601

[0103] The therapeutic intervention may comprise recommending the subject for a secondary clinical test to confirm a diagnosis of the bladder cancer. This secondary clinical test may comprise an imaging test, a blood test, a computed tomography (CT) scan, a magnetic resonance imaging (MRI) scan, an ultrasound scan, a chest X-ray, a positron emission tomography (PET) scan, a PET-CT scan, a biological cytology, or any combination thereof.

[0104] The quantitative measures of sequence reads of the dataset at the panel of bladder cancer-associated genomic loci (e.g., quantitative measures of RNA transcripts or DNA at the bladder cancer-associated genomic loci), proteomic data comprising quantitative measures of proteins of the dataset at a panel of bladder cancer-associated proteins, and / or metabolome data comprising quantitative measures of a panel of bladder cancer-associated metabolites may be assessed over a duration of time to monitor a patient (e.g., subject who has bladder cancer or who is being treated for bladder cancer). In such cases, the quantitative measures of the dataset of the patient may change during the course of treatment. For example, the quantitative measures of the dataset of a patient with decreasing risk of the bladder cancer due to an effective treatment may shift toward the profile or distribution of a healthy subject (e.g., a subject without a bladder cancer complication). Conversely, for example, the quantitative measures of the dataset of a patient with increasing risk of the bladder cancer due to an ineffective treatment may shift toward the profile or distribution of a subject with higher risk of the bladder cancer or a more advanced bladder cancer.

[0105] The bladder cancer of the subject may be monitored by monitoring a course of treatment for treating the bladder cancer of the subject. The monitoring may comprise assessing the bladder cancer of the subject at two or more time points. The assessing may be based at least on the quantitative measures of sequence reads of the dataset at a panel of bladder cancer-associated genomic loci (e.g., quantitative measures of RNA transcripts or DNA at the bladder cancer-associated genomic loci), proteomic data comprising quantitative measures of proteins of the dataset at a panel of bladder cancer-associated proteins, and / or metabolome data comprising quantitative measures of a panel of bladder cancer-associated metabolites determined at each of the two or more time points.

[0106] In some embodiments, a difference in the quantitative measures of sequence reads of the dataset at a panel of bladder cancer-associated genomic loci (e.g., quantitative measures of RNA transcripts or DNA at the bladder cancer-associated genomic loci), proteomic data comprising quantitative measures of proteins of the dataset at a panel of bladder cancer-associated proteins, and / or metabolome data comprising quantitative measures of a panel of bladder cancer-associated metabolites determined between the two or more time points may be indicative of one or more clinical indications, such as (i) a diagnosis of the bladder cancer of the subject, (ii) aAttomey Docket No. 70528-701.601prognosis of the bladder cancer of the subject, (iii) an increased risk of the bladder cancer of the subject, (iv) a decreased risk of the bladder cancer of the subject, (v) an efficacy of the course of treatment for treating the bladder cancer of the subject, and (vi) a non-efficacy of the course of treatment for treating the bladder cancer of the subject.

[0107] In some embodiments, a difference in the quantitative measures of sequence reads of the dataset at a panel of bladder cancer-associated genomic loci (e.g., quantitative measures of RNA transcripts or DNA at the bladder cancer-associated genomic loci), proteomic data comprising quantitative measures of proteins of the dataset at a panel of bladder cancer-associated proteins, and / or metabolome data comprising quantitative measures of a panel of bladder cancer-associated metabolites determined between the two or more time points may be indicative of a diagnosis of the bladder cancer of the subject. For example, if the bladder cancer was not detected in the subject at an earlier time point but was detected in the subject at a later time point, then the difference is indicative of a diagnosis of the bladder cancer of the subject. A clinical action or decision may be made based on this indication of diagnosis of the bladder cancer of the subject, such as, for example, prescribing a new therapeutic intervention for the subject. The clinical action or decision may comprise recommending the subject for a secondary clinical test to confirm the diagnosis of the bladder cancer. This secondary clinical test may comprise an imaging test, a blood test, a computed tomography (CT) scan, a magnetic resonance imaging (MRI) scan, an ultrasound scan, a chest X-ray, a positron emission tomography (PET) scan, a PET-CT scan, a biological cytology, or any combination thereof.

[0108] In some embodiments, a difference in the quantitative measures of sequence reads of the dataset at a panel of bladder cancer-associated genomic loci (e.g., quantitative measures of RNA transcripts or DNA at the bladder cancer-associated genomic loci), proteomic data comprising quantitative measures of proteins of the dataset at a panel of bladder cancer-associated proteins, and / or metabolome data comprising quantitative measures of a panel of bladder cancer-associated metabolites determined between the two or more time points may be indicative of a prognosis of the bladder cancer of the subject.

[0109] In some embodiments, a difference in the quantitative measures of sequence reads of the dataset at a panel of bladder cancer-associated genomic loci (e.g., quantitative measures of RNA transcripts or DNA at the bladder cancer-associated genomic loci), proteomic data comprising quantitative measures of proteins of the dataset at a panel of bladder cancer-associated proteins, and / or metabolome data comprising quantitative measures of a panel of bladder cancer-associated metabolites determined between the two or more time points may be indicative of the subject having an increased risk of the bladder cancer. For example, if the bladder cancer was detected in the subject both at an earlier time point and at a later time point, and if the differenceAttomey Docket No. 70528-701.601is a positive difference (e.g., the quantitative measures of sequence reads of the dataset at a panel of bladder cancer-associated genomic loci (e.g., quantitative measures of RNA transcripts or DNA at the bladder cancer-associated genomic loci), proteomic data comprising quantitative measures of proteins of the dataset at a panel of bladder cancer-associated proteins, and / or metabolome data comprising quantitative measures of a panel of bladder cancer-associated metabolites increased from the earlier time point to the later time point), then the difference may be indicative of the subject having an increased risk of the bladder cancer. A clinical action or decision may be made based on this indication of the increased risk of the bladder cancer, e.g., prescribing a new therapeutic intervention or switching therapeutic interventions (e.g., ending a current treatment and prescribing a new treatment) for the subject. The clinical action or decision may comprise recommending the subject for a secondary clinical test to confirm the increased risk of the bladder cancer. This secondary clinical test may comprise an imaging test, a blood test, a computed tomography (CT) scan, a magnetic resonance imaging (MRI) scan, an ultrasound scan, a chest X-ray, a positron emission tomography (PET) scan, a PET-CT scan, a biological cytology, or any combination thereof.

[0110] In some embodiments, a difference in the quantitative measures of sequence reads of the dataset at a panel of bladder cancer-associated genomic loci (e.g., quantitative measures of RNA transcripts or DNA at the bladder cancer-associated genomic loci), proteomic data comprising quantitative measures of proteins of the dataset at a panel of bladder cancer-associated proteins, and / or metabolome data comprising quantitative measures of a panel of bladder cancer-associated metabolites determined between the two or more time points may be indicative of the subject having a decreased risk of the bladder cancer. For example, if the bladder cancer was detected in the subject both at an earlier time point and at a later time point, and if the difference is a negative difference (e.g., the quantitative measures of sequence reads of the dataset at a panel of bladder cancer-associated genomic loci (e.g., quantitative measures of RNA transcripts or DNA at the bladder cancer-associated genomic loci), proteomic data comprising quantitative measures of proteins of the dataset at a panel of bladder cancer-associated proteins, and / or metabolome data comprising quantitative measures of a panel of bladder cancer-associated metabolites decreased from the earlier time point to the later time point), then the difference may be indicative of the subject having a decreased risk of the bladder cancer. A clinical action or decision may be made based on this indication of the decreased risk of the bladder cancer (e.g., continuing or ending a current therapeutic intervention) for the subject. The clinical action or decision may comprise recommending the subject for a secondary clinical test to confirm the decreased risk of the bladder cancer. This secondary clinical test may comprise an imaging test, a blood test, a computed tomography (CT) scan, a magnetic resonance imaging (MRI) scan, anAttomey Docket No. 70528-701.601ultrasound scan, a chest X-ray, a positron emission tomography (PET) scan, a PET-CT scan, a biological cytology, or any combination thereof.[oni] In some embodiments, a difference in the quantitative measures of sequence reads of the dataset at a panel of bladder cancer-associated genomic loci (e.g., quantitative measures of RNA transcripts or DNA at the bladder cancer-associated genomic loci), proteomic data comprising quantitative measures of proteins of the dataset at a panel of bladder cancer-associated proteins, and / or metabolome data comprising quantitative measures of a panel of bladder cancer-associated metabolites determined between the two or more time points may be indicative of an efficacy of the course of treatment for treating the bladder cancer of the subject. For example, if the bladder cancer was detected in the subject at an earlier time point but was not detected in the subject at a later time point, then the difference may be indicative of an efficacy of the course of treatment for treating the bladder cancer of the subject. A clinical action or decision may be made based on this indication of the efficacy of the course of treatment for treating the bladder cancer of the subject, e.g., continuing or ending a current therapeutic intervention for the subject. The clinical action or decision may comprise recommending the subject for a secondary clinical test to confirm the efficacy of the course of treatment for treating the bladder cancer. This secondary clinical test may comprise an imaging test, a blood test, a computed tomography (CT) scan, a magnetic resonance imaging (MRI) scan, an ultrasound scan, a chest X-ray, a positron emission tomography (PET) scan, a PET-CT scan, a biological cytology, or any combination thereof.

[0112] In some embodiments, a difference in the quantitative measures of sequence reads of the dataset at a panel of bladder cancer-associated genomic loci (e.g., quantitative measures of RNA transcripts or DNA at the bladder cancer-associated genomic loci), proteomic data comprising quantitative measures of proteins of the dataset at a panel of bladder cancer-associated proteins, and / or metabolome data comprising quantitative measures of a panel of bladder cancer-associated metabolites determined between the two or more time points may be indicative of a non-efficacy of the course of treatment for treating the bladder cancer of the subject. For example, if the bladder cancer was detected in the subject both at an earlier time point and at a later time point, and if the difference is a positive or zero difference (e.g., the quantitative measures of sequence reads of the dataset at a panel of bladder cancer-associated genomic loci (e.g., quantitative measures of RNA transcripts or DNA at the bladder cancer-associated genomic loci), proteomic data comprising quantitative measures of proteins of the dataset at a panel of bladder cancer-associated proteins, and / or metabolome data comprising quantitative measures of a panel of bladder cancer-associated metabolites increased or remained at a constant level from the earlier time point to the later time point), and if an efficacious treatment wasAttomey Docket No. 70528-701.601indicated at an earlier time point, then the difference may be indicative of a non-efficacy of the course of treatment for treating the bladder cancer of the subject. A clinical action or decision may be made based on this indication of the non-efficacy of the course of treatment for treating the bladder cancer of the subject, e.g., ending a current therapeutic intervention and / or switching to (e.g., prescribing) a different new therapeutic intervention for the subject. The clinical action or decision may comprise recommending the subject for a secondary clinical test to confirm the non-efficacy of the course of treatment for treating the bladder cancer. This secondary clinical test may comprise an imaging test, a blood test, a computed tomography (CT) scan, a magnetic resonance imaging (MRI) scan, an ultrasound scan, a chest X-ray, a positron emission tomography (PET) scan, a PET-CT scan, a biological cytology, or any combination thereof.

[0113] In some embodiments, the trained algorithm may process clinical health data of the subject. The clinical health data may comprise one or more quantitative measures of the subject, such as age, weight, height, body mass index (BMI), blood pressure, heart rate, and glucose levels. As another example, the clinical health data can comprise one or more categorical measures, such as race, ethnicity, history of medication or other clinical treatment, history of tobacco use, history of alcohol consumption, daily activity or fitness level, genetic test results, blood test results, and imaging results.

[0114] In some embodiments, the computer-implemented methods described herein are performed using a computer or mobile device application. For example, a subject can use a computer or mobile device application to input her own clinical health data, including quantitative and / or categorical measures. The computer or mobile device application can then use a trained algorithm to process the clinical health data to determine a risk score indicative of the risk of bladder cancer of the subject. The computer or mobile device application can then display a report indicative of the risk score indicative of the risk of bladder cancer of the subject.

[0115] In some embodiments, the risk score indicative of the risk of bladder cancer of the subject can be refined by performing one or more subsequent clinical tests for the subject. For example, the subject can be referred by a physician for one or more subsequent clinical tests (e.g., an imaging or a blood test) based on the initial risk score. Next, the computer or mobile device application may process results from the one or more subsequent clinical tests using a trained algorithm to determine an updated risk score indicative of the risk of bladder cancer of the subject.

[0116] After the bladder cancer is identified or an increased risk of the bladder cancer is monitored in the subject, a report may be electronically outputted that is indicative of (e.g., identifies or provides an indication of) the bladder cancer of the subject. The subject may not display a bladder cancer (e.g., is asymptomatic of the bladder cancer). The report may beAttomey Docket No. 70528-701.601presented on a graphical user interface (GUI) of an electronic device of a user. The user may be the subject, a caretaker, a physician, a nurse, or another health care worker.

[0117] The report may include one or more clinical indications such as (i) a diagnosis of the bladder cancer of the subject, (ii) a prognosis of the bladder cancer of the subject, (iii) an increased risk of the bladder cancer of the subject, (iv) a decreased risk of the bladder cancer of the subject, (v) an efficacy of the course of treatment for treating the bladder cancer of the subject, and (vi) a non-efficacy of the course of treatment for treating the bladder cancer of the subject. The report may include one or more clinical actions or decisions made based on these one or more clinical indications. Such clinical actions or decisions may be directed to therapeutic interventions or further clinical assessment or testing of the bladder cancer of the subject.

[0118] For example, a clinical indication of a diagnosis of the bladder cancer of the subject may be accompanied with a clinical action of prescribing a new therapeutic intervention for the subject. As another example, a clinical indication of an increased risk of the bladder cancer of the subject may be accompanied with a clinical action of prescribing a new therapeutic intervention or switching therapeutic interventions (e.g., ending a current treatment and prescribing a new treatment) for the subject. As another example, a clinical indication of a decreased risk of the bladder cancer of the subject may be accompanied with a clinical action of continuing or ending a current therapeutic intervention for the subject. As another example, a clinical indication of an efficacy of the course of treatment for treating the bladder cancer of the subject may be accompanied with a clinical action of continuing or ending a current therapeutic intervention for the subject. As another example, a clinical indication of a non-efficacy of the course of treatment for treating the bladder cancer of the subject may be accompanied with a clinical action of ending a current therapeutic intervention and / or switching to (e.g., prescribing) a different new therapeutic intervention for the subject.

[0119] Disclosed herein are computer-implemented methods comprising receiving the genetic data and the transcriptomic data for a subject and analyzing the genetic data and the transcriptomic data using a trained machine learning algorithm to output the degree of cancer risk. In some instances, the output may be indicative of a degree of cancer risk of the subject. In some instances, the degree of cancer risk is a categorical cancer risk selected from among a plurality of distinct categorical cancer risks. In some instances, the distinct categorical cancer risks comprises a high cancer risk, an intermediate cancer risk, or a low cancer risk, or any combination thereof. In some embodiments, the method further comprises calculating a risk score corresponding to the high cancer risk, wherein the risk score exceeds the threshold of 0.54. In some embodiments, the method further comprises calculating a risk score corresponding toAttomey Docket No. 70528-701.601the high cancer risk, wherein the risk score is from 0.15 to 0.54. In some embodiments, the method further comprises calculating a risk score corresponding to the high cancer risk, wherein the risk score is below 0.15. The trained machine learning algorithm may determine an assigned classification. In some instances, the assigned classification is a positive cancer risk. The positive cancer risk classification may be output when the distinct categorical cancer risk of a subject comprises a high cancer risk or the intermediate cancer risk. In some instances, the assigned classification is a negative cancer risk. The negative cancer risk classification may be output when the distinct categorical cancer risk of a subject comprises a low cancer risk. In some embodiments, the method may further comprise generating an electronic report comprising the output indicative of the subject as having the degree of cancer risk.

[0120] Provided herein is a trained machine learning classifier. In some instances, the trained machine learning classifier may be indicative of a degree of cancer risk of the subject. In some instances, the degree of cancer risk is a categorical cancer risk selected from among a plurality of distinct categorical cancer risks. In some instances, the distinct categorical cancer risks comprises a high cancer risk, an intermediate cancer risk, or a low cancer risk, or any combination thereof. In some embodiments, the method further comprises calculating a risk score corresponding to the high cancer risk, wherein the risk score exceeds the threshold of 0.54. In some embodiments, the method further comprises calculating a risk score corresponding to the high cancer risk, wherein the risk score is from 0.15 to 0.54. In some embodiments, the method further comprises calculating a risk score corresponding to the high cancer risk, wherein the risk score is below 0.15. The trained machine learning algorithm may determine an assigned classification. In some instances, the assigned classification is a positive cancer risk. The positive cancer risk classification may be output when the distinct categorical cancer risk of a subject comprises a high cancer risk or the intermediate cancer risk. In some instances, the assigned classification is a negative cancer risk. The negative cancer risk classification may be output when the distinct categorical cancer risk of a subject comprises a low cancer risk. In some embodiments, the method may further comprise generating an electronic report comprising the output indicative of the subject as having the degree of cancer risk.

[0121] Disclosed herein, in some embodiments, are methods comprising generating an electronic report comprising the output of the trained machine learning algorithm indicative of the subject as having the degree of cancer risk based, at least in part, on the genetic data and transcriptomic data derived from a sample obtained from the subject. In some embodiments, the electronic report comprises the risk score. In some embodiments, the electronic report comprises the assigned classification. In some instances, the assigned classification is a positive cancer risk. The positive cancer risk classification may be output when the distinct categorical cancerAttomey Docket No. 70528-701.601risk of a subject comprises a high cancer risk or the intermediate cancer risk. In some instances, the assigned classification is a negative cancer risk. The negative cancer risk classification may be output when the distinct categorical cancer risk of a subject comprises a low cancer risk. II. SYSTEMS

[0122] Disclosed herein are systems for analyzing a sample. In some embodiments, the system comprises a test kit configured to collect and prepare a urine sample. In some embodiments, the system comprises a PCR machine. In some embodiments, a first data set (e.g., genetic data) and a second data set (e.g., transcriptomic data) using a trained machine learning algorithm to determine an output indicative of a degree of cancer risk. In some embodiments, the first data set comprises genetic data corresponding to one or more polymorphic loci or one or more genotypes at the polymorphic loci. In some embodiments, the second data set comprises transcriptomic data corresponding to an amount or a level of one or more transcriptomic markers.

[0123] Disclosed herein, in some embodiments, are systems comprising a genotyping assay. An assaying device that may be useful for the systems herein include a quantitative polymerase chain reaction (qPCR), nucleic acid sequencing reaction, or a genotyping array, polymerase chain reactions (PCRs), or droplet-digital PCRs (ddPCR). In some embodiments, the ddPCR utilizes a Taq polymerase in a standard PCR reaction to amplify a target DNA fragment from a sample. ddPCR may also partition the PCR reaction into thousands of individual reaction vessels before an amplification step. ddPCR may also comprise acquiring data at a reaction end point. In some embodiments, the genotyping assay detects the genotype at one or more polymorphisms. In some embodiments, the genotyping assay detects the genotype at one or more polymorphisms generates a first data set. In some embodiments, the genotypes are heterozygous for a risk allele at the one or more polymorphisms. In some instances, the genotyping assay comprises a polymerase chain reaction (PCR). In some embodiments, the PCR is a quantitative reversetranscription PCR (rt-PCR). In some embodiments, the one or more polymorphisms comprise rsl21913482, rsl21913483, rsl21913479, rsl21913485, rsl242535815, rsl561215364, or a polymorphism in linkage disequilibrium therewith as determined by a coefficient of determination R2of at least 0.85, or any combination thereof.

[0124] The genetic data may include genotypes at one or more polymorphic loci. In some embodiments, the one or more polymorphic loci comprises rsl21913482, rsl21913483, rsl21913479, rsl21913485, rsl242535815, rsl561215364, or a polymorphism in linkage disequilibrium therewith as determined by a coefficient of determination R2of at least 0.85, or any combination thereof. Linkage disequilibrium refers to the non-random association of alleles or indels in different gene loci in a given population. LD may be measured by a D’ valueAttomey Docket No. 70528-701.601corresponding to the difference between an observed and expected allele or indel frequencies in the population (D=Pab-PaPb), which is scaled by a theoretical maximum value of D. LD may be defined by an R2value corresponding to the difference between an observed and expected unit of risk frequencies in the population (D=Pab-PaPb), which is scaled by the individual frequencies of the different loci. In some embodiments, the linkage disequilibrium may be defined by an R2value of at least 0.80, 0.85, 0.90, 0.95, or 1.0. In some embodiments, the one or more polymorphic loci comprises two or more of rsl21913482, rsl21913483, rsl21913479, rsl21913485, rsl242535815, rsl561215364, or a polymorphism in linkage disequilibrium therewith as determined by a coefficient of determination R2of at least 0.85, or any combination thereof. In some embodiments, the one or more polymorphic loci comprises three or more of rsl21913482, rsl21913483, rsl21913479, rsl21913485, rsl242535815, rsl561215364, or a polymorphism in linkage disequilibrium therewith as determined by a coefficient of determination R2of at least 0.85, or any combination thereof. In some embodiments, the one or more polymorphic loci comprises four or more of rsl21913482, rsl21913483, rsl21913479, rsl21913485, rsl242535815, rsl561215364, or a polymorphism in linkage disequilibrium therewith as determined by a coefficient of determination R2of at least 0.85, or any combination thereof. In some embodiments, the one or more polymorphic loci comprises five or more of rsl21913482, rsl21913483, rsl21913479, rsl21913485, rsl242535815, rsl561215364, or a polymorphism in linkage disequilibrium therewith as determined by a coefficient of determination R2of at least 0.85, or any combination thereof. In some embodiments, the one or more polymorphic loci comprises six or more of rsl21913482, rsl21913483, rsl21913479, rsl21913485, rsl242535815, rsl561215364, or a polymorphism in linkage disequilibrium therewith as determined by a coefficient of determination R2of at least 0.85, or any combination thereof. In some embodiments, the one or more polymorphisms comprise rsl21913482, rsl21913483, rsl21913479, rsl21913485, rsl242535815, and rsl561215364.

[0125] In some embodiments, the one or more polymorphisms comprise rsl21913482. In some embodiments, rsl21913482 is referred to as MT742 rsl21913482. rsl21913482 is a single nucleotide variant of the fibroblast growth factor receptor 3 (FGFR3) gene. rsl21913482 variations are associated with UC. In some embodiments, the position of rsl21913482 is at chr4: 1801837 (GRCh38.pl4). In some embodiments, the position of rsl21913482 is at chr4: 1803564 (GRCh37). Effects of rsl21913482 variations may include a protein change of R248C.

[0126] In some embodiments, the one or more polymorphisms comprise rsl21913483. In some embodiments, rsl21913483 is referred to as MT746 rsl21913483. rsl21913483 is a single nucleotide variant of the fibroblast growth factor receptor 3 (FGFR3) gene. rsl21913483Attomey Docket No. 70528-701.601variations are associated with UC. In some embodiments, the position of rsl21913483 is at chr4: 1801841 (GRCh38.pl4). In some embodiments, the position of rsl21913483 is at chr4: 1803568 (GRCh37). Effects of rsl21913482 variations may include a protein change of S249F or S249C.

[0127] In some embodiments, the one or more polymorphisms comprise rsl21913479. In some embodiments, rsl21913479 is referred to as MT1114 rsl21913479. rsl21913479 is a single nucleotide variant of the fibroblast growth factor receptor 3 (FGFR3) gene. rsl21913479 variations are associated with UC. In some embodiments, the position of rsl21913479 is at chr4: 1804362 (GRCh38.pl4). In some embodiments, the position of rsl21913479 is at chr4: 1806089 (GRCh37). Effects of rsl21913479 variations may include a protein change of G370S orG370C.

[0128] In some embodiments, the one or more polymorphisms comprise rsl21913485. In some embodiments, rsl21913485 is referred to as MT1124 rsl21913485. rsl21913485 is a single nucleotide variant of the fibroblast growth factor receptor 3 (FGFR3) gene. rsl21913485 variations are associated with UC. In some embodiments, the position of rsl21913485 is at chr4: 1804372 (GRCh38.pl4). In some embodiments, the position of rsl21913485 is at chr4: 1806099 (GRCh37). Effects of rsl21913485 variations may include a protein change of Y373C or Y375C.

[0129] In some embodiments, the one or more polymorphisms comprise rsl242535815. In some embodiments, rsl242535815 is referred to as MT124 rsl242535815. rsl242535815 is a single nucleotide variant of the telomerase reverse transcriptase (TERT) gene in the promotor region. rsl242535815 variations are associated with UC. In some embodiments, the position of rsl242535815 is at chr5:1295113 (GRCh38.pl4). In some embodiments, the position of rsl242535815 is at chr5: 1295228 (GRCh37). Effects of rsl242535815 c.-124C>T mutations may include increases in expression of TERT.

[0130] In some embodiments, the one or more polymorphisms comprise rsl561215364. In some embodiments, rsl561215364 is referred to as MT146 rsl561215364. rsl561215364 is a single nucleotide variant of the telomerase reverse transcriptase (TERT) gene in the promotor region. rsl561215364 variations are associated with UC. In some embodiments, the position of rsl561215364 is at chr5:1295135 (GRCh38.pl4). In some embodiments, the position of rsl561215364 is at chr5: 1295250 (GRCh37). Effects of rsl561215364 c.-146C>T mutations may induce over expression of proteins.

[0131] In some embodiments, the polymorphic loci include one or polymorphic loci in linkage disequilibrium therewith. Linkage disequilibrium may be defined by an r2value of at least 0.80, 0.85, 0.90, 0.95, or 1.0.Attomey Docket No. 70528-701.601

[0132] In some embodiments, polymorphic locus or a genotype at a polymorphic locus is associated with the cancer (e.g., bladder cancer) with a P value of at most about 1.0 x 10'6, about 1.0 x 10'7, about 1.0 x 10'8, about 1.0 x 10'9, about 1.0 x 10'10, about 1.0 x IO'20, about 1.0 x 10"30, about 1.0 x IO'40, about 1.0 x 10'50, about 1.0 x IO'60, about 1.0 x IO'70, about 1.0 x IO'80, about 1.0 x IO'90, or about 1.0 x IO'100.

[0133] Disclosed herein, in some embodiments, are systems comprising a transcriptomic assay. An assaying device that may be useful for the methods herein include a polymerase chain reaction (PCR), a quantitative polymerase chain reaction (qPCR), a quantitative reversetranscription PCR (rtPCR), nucleic acid sequencing reaction, a genotyping array, or dropletdigital PCRs (ddPCR). In some embodiments, the assay device comprises a quantitative rtPCR. In some embodiments, the quantitative rtPCR is performed in one step, wherein a reverse transcription step and a PCR step are performed in a single tube and buffer, using a reverse transcriptase along with a DNA polymerase, wherein only sequence-specific primers are utilized. In some embodiments, the quantitative rtPCR is performed in two steps, wherein the reverse transcription step and the PCR step are performed in separate tubes comprising different optimized buffers, reaction conditions, and priming strategies. The transcriptomic data may include an amount or a level of one or more transcriptomic markers. In some embodiments, the one or more transcriptomic markers comprise Midkine (MDK), Cyclin Dependent Kinase 1 (CDK1), Insulin Like Growth Factor Binding Protein 5 (IGFBP5), Homeobox A13 (HOXA13), or C-X-C Motif Chemokine Receptor 2 (CXCR2), or any combination thereof. In some embodiments, the one or more transcriptomic markers comprises two or more of Midkine (MDK), Cyclin Dependent Kinase 1 (CDK1), Insulin Like Growth Factor Binding Protein 5 (IGFBP5), Homeobox Al 3 (HOXA13), or C-X-C Motif Chemokine Receptor 2 (CXCR2), or any combination thereof. In some embodiments, the one or more transcriptomic markers comprises three or more of Midkine (MDK), Cyclin Dependent Kinase 1 (CDK1), Insulin Like Growth Factor Binding Protein 5 (IGFBP5), Homeobox A13 (HOXA13), or C-X-C Motif Chemokine Receptor 2 (CXCR2), or any combination thereof. In some embodiments, the one or more transcriptomic markers comprises four or more of Midkine (MDK), Cyclin Dependent Kinase 1 (CDK1), Insulin Like Growth Factor Binding Protein 5 (IGFBP5), Homeobox A13 (HOXA13), or C-X-C Motif Chemokine Receptor 2 (CXCR2), or any combination thereof. In some embodiments, the one or more transcriptomic markers comprise Midkine (MDK), Cyclin Dependent Kinase 1 (CDK1), Insulin Like Growth Factor Binding Protein 5 (IGFBP5), Homeobox Al 3 (HOXA13), and C-X-C Motif Chemokine Receptor 2 (CXCR2).

[0134] In some embodiments, the one or more transcriptomic markers comprise Midkine (MDK). MDK is a multifunctional protein that forms a family of heparin-binding growthAttomey Docket No. 70528-701.601factors. MDK may be overexpressed in many types of cancer, including bladder cancer and UC. Overexpression of MDK may be especially noticeable during progression of a tumor into more advanced stages. MDK expression in tumors may be determined by multiple forms of analysis, including blood and urine analysis.

[0135] In some embodiments, the one or more transcriptomic markers comprise Cyclin Dependent Kinase 1 (CDK1). CDK1 is kinase that is crucial in cell cycle progression. One indicators of malignancy, such as unrestricted cell proliferation, may be caused by alterations in CDK1 activity.

[0136] In some embodiments, the one or more transcriptomic markers comprise Insulin Like Growth Factor Binding Protein 5 (IGFBP5). IGFBP5 is associated with several cancers, including bladder cancer and UC.

[0137] In some embodiments, the one or more transcriptomic markers comprise Homeobox Al 3 (HOXA13). HOXA13 is associated with several cancers, including bladder cancer and UC.

[0138] In some embodiments, the one or more transcriptomic markers comprise C-X-C Motif Chemokine Receptor 2 (CXCR2). CXCR2 is an inflammation marker.

[0139] In some embodiments, the amount of the transcriptomic marker is an absolute amount, whereas a level of the transcriptomic marker may be an expression level that is relative to a reference expression level for the transcriptomic marker.

[0140] Disclosed herein, in some embodiments, are kits of the present disclosure, such as sample collection and / or preparation kits. In other embodiments, the kits contain all of the components necessary and / or sufficient to perform an analysis of a urine sample, including all controls, directions for performing assays, and any necessary software for analysis and presentation of results.A. Computing system

[0141] Referring to FIG. 1, a block diagram is shown depicting an exemplary machine that includes a computer system 100 (e.g., a processing or computing system) within which a set of instructions can execute for causing a device to perform or execute any one or more of the aspects and / or methodologies for static code scheduling of the present disclosure. The components in FIG. 1 are examples only and do not limit the scope of use or functionality of any hardware, software, embedded logic component, or a combination of two or more such components implementing particular embodiments.

[0142] Computer system 100 may include one or more processors 101, a memory 103, and a storage 108 that communicate with each other, and with other components, via a bus 140. The bus 140 may also link a display 132, one or more input devices 133 (which may, for example,Attomey Docket No. 70528-701.601include a keypad, a keyboard, a mouse, a stylus, etc.), one or more output devices 134, one or more storage devices 135, and various tangible storage media 136. All of these elements may interface directly or via one or more interfaces or adaptors to the bus 140. For instance, the various tangible storage media 136 can interface with the bus 140 via storage medium interface 126. Computer system 100 may have any suitable physical form, including but not limited to one or more integrated circuits (ICs), printed circuit boards (PCBs), mobile handheld devices (such as mobile telephones or PDAs), laptop or notebook computers, distributed computer systems, computing grids, or servers.

[0143] Computer system 100 includes one or more processor(s) 101 (e.g., central processing units (CPUs), general purpose graphics processing units (GPGPUs), or quantum processing units (QPUs)) that carry out functions. Processor(s) 101 optionally contains a cache memory unit 102 for temporary local storage of instructions, data, or computer addresses. Processor(s) 101 are configured to assist in execution of computer readable instructions. Computer system 100 may provide functionality for the components depicted in FIG. 1 as a result of the processor(s) 101 executing non-transitory, processor-executable instructions embodied in one or more tangible computer-readable storage media, such as memory 103, storage 108, storage devices 135, and / or storage medium 136. The computer-readable media may store software that implements particular embodiments, and processor(s) 101 may execute the software. Memory 103 may read the software from one or more other computer-readable media (such as mass storage device(s) 135, 136) or from one or more other sources through a suitable interface, such as network interface 120. The software may cause processor(s) 101 to carry out one or more processes or one or more steps of one or more processes described or illustrated herein. Carrying out such processes or steps may include defining data structures stored in memory 103 and modifying the data structures as directed by the software.

[0144] The memory 103 may include various components (e.g., machine readable media) including, but not limited to, a random access memory component (e.g., RAM 104) (e.g., static RAM (SRAM), dynamic RAM (DRAM), ferroelectric random access memory (FRAM), phasechange random access memory (PRAM), etc.), a read-only memory component (e.g., ROM 105), and any combinations thereof. ROM 105 may act to communicate data and instructions unidirectionally to processor(s) 101, and RAM 104 may act to communicate data and instructions bidirectionally with processor(s) 101. ROM 105 and RAM 104 may include any suitable tangible computer-readable media described below. In one example, a basic input / output system 106 (BIOS), including basic routines that help to transfer information between elements within computer system 100, such as during start-up, may be stored in the memory 103.Attomey Docket No. 70528-701.601

[0145] Fixed storage 108 is connected bidirectionally to processor(s) 101, optionally through storage control unit 107. Fixed storage 108 provides additional data storage capacity and may also include any suitable tangible computer-readable media described herein. Storage 108 may be used to store operating system 109, executable(s) 110, data 111, applications 112 (application programs), and the like. Storage 108 can also include an optical disk drive, a solid-state memory device (e.g., flash-based systems), or a combination of any of the above. Information in storage 108 may, in appropriate cases, be incorporated as virtual memory in memory 103.

[0146] In one example, storage device(s) 135 may be removably interfaced with computer system 100 (e.g., via an external port connector (not shown)) via a storage device interface 125.Particularly, storage device(s) 135 and an associated machine-readable medium may provide non-volatile and / or volatile storage of machine-readable instructions, data structures, program modules, and / or other data for the computer system 100. In one example, software may reside, completely or partially, within a machine-readable medium on storage device(s) 135. In another example, software may reside, completely or partially, within processor(s) 101.

[0147] Bus 140 connects a wide variety of subsystems. Herein, reference to a bus may encompass one or more digital signal lines serving a common function, where appropriate. Bus 140 may be any of several types of bus structures including, but not limited to, a memory bus, a memory controller, a peripheral bus, a local bus, and any combinations thereof, using any of a variety of bus architectures. As an example and not by way of limitation, such architectures include an Industry Standard Architecture (ISA) bus, an Enhanced ISA (EISA) bus, a Micro Channel Architecture (MCA) bus, a Video Electronics Standards Association local bus (VLB), a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCLX) bus, an Accelerated Graphics Port (AGP) bus, HyperTransport (HTX) bus, serial advanced technology attachment (SATA) bus, and any combinations thereof.

[0148] Computer system 100 may also include an input device 133. In one example, a user of computer system 100 may enter commands and / or other information into computer system 100 via input device(s) 133. Examples of an input device(s) 133 include, but are not limited to, an alpha-numeric input device (e.g., a keyboard), a pointing device (e.g., a mouse or touchpad), a touchpad, a touch screen, a multi-touch screen, a joystick, a stylus, a gamepad, an audio input device (e.g., a microphone, a voice response system, etc.), an optical scanner, a video or still image capture device (e.g., a camera), and any combinations thereof. In some embodiments, the input device is a Kinect, Leap Motion, or the like. Input device(s) 133 may be interfaced to bus 140 via any of a variety of input interfaces 123 (e.g., input interface 123) including, but not limited to, serial, parallel, game port, USB, FIREWIRE, THUNDERBOLT, or any combination of the above.Attomey Docket No. 70528-701.601

[0149] In particular embodiments, when computer system 100 is connected to network 130, computer system 100 may communicate with other devices, specifically mobile devices and enterprise systems, distributed computing systems, cloud storage systems, cloud computing systems, and the like, connected to network 130. Communications to and from computer system 100 may be sent through network interface 120. For example, network interface 120 may receive incoming communications (such as requests or responses from other devices) in the form of one or more packets (such as Internet Protocol (IP) packets) from network 130, and computer system 100 may store the incoming communications in memory 103 for processing. Computer system 100 may similarly store outgoing communications (such as requests or responses to other devices) in the form of one or more packets in memory 103 and communicated to network 130 from network interface 120. Processor(s) 101 may access these communication packets stored in memory 103 for processing.

[0150] Examples of the network interface 120 include, but are not limited to, a network interface card, a modem, and any combination thereof. Examples of a network 130 or network segment 130 include, but are not limited to, a distributed computing system, a cloud computing system, a wide area network (WAN) (e.g., the Internet, an enterprise network), a local area network (LAN) (e.g., a network associated with an office, a building, a campus or other relatively small geographic space), a telephone network, a direct connection between two computing devices, a peer-to-peer network, and any combinations thereof. A network, such as network 130, may employ a wired and / or a wireless mode of communication. In general, any network topology may be used.

[0151] Information and data can be displayed through a display 132. Examples of a display 132 include, but are not limited to, a cathode ray tube (CRT), a liquid crystal display (LCD), a thin film transistor liquid crystal display (TFT-LCD), an organic liquid crystal display (OLED) such as a passive-matrix OLED (PMOLED) or active-matrix OLED (AMOLED) display, a plasma display, and any combinations thereof. The display 132 can interface to the processor(s) 101, memory 103, and fixed storage 108, as well as other devices, such as input device(s) 133, via the bus 140. The display 132 is linked to the bus 140 via a video interface 122, and transport of data between the display 132 and the bus 140 can be controlled via the graphics control 121. In some embodiments, the display is a video projector. In some embodiments, the display is a headmounted display (HMD) such as a VR headset. In further embodiments, suitable VR headsets include, by way of non-limiting examples, HTC Vive, Oculus Rift, Samsung Gear VR, Microsoft HoloLens, Razer OSVR, FOVE VR, Zeiss VR One, Avegant Glyph, Freefly VR headset, and the like. In still further embodiments, the display is a combination of devices such as those disclosed herein.Attomey Docket No. 70528-701.601

[0152] In addition to a display 132, computer system 100 may include one or more other peripheral output devices 134 including, but not limited to, an audio speaker, a printer, a storage device, and any combinations thereof. Such peripheral output devices may be connected to the bus 140 via an output interface 124. Examples of an output interface 124 include, but are not limited to, a serial port, a parallel connection, a USB port, a FIREWIRE port, a THUNDERBOLT port, and any combinations thereof.

[0153] In addition or as an alternative, computer system 100 may provide functionality as a result of logic hardwired or otherwise embodied in a circuit, which may operate in place of or together with software to execute one or more processes or one or more steps of one or more processes described or illustrated herein. Reference to software in this disclosure may encompass logic, and reference to logic may encompass software. Moreover, reference to a computer-readable medium may encompass a circuit (such as an IC) storing software for execution, a circuit embodying logic for execution, or both, where appropriate. The present disclosure encompasses any suitable combination of hardware, software, or both.

[0154] Various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and operations have been described above generally in terms of their functionality.

[0155] The various illustrative logical blocks, modules, and circuits described in connection with the embodiments disclosed herein may be implemented or performed with a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.

[0156] The operation of a method or algorithm described in connection with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by one or more processor(s), or in a combination of the two. A software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium. An exemplary storageAttomey Docket No. 70528-701.601medium is coupled to the processor such the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In the alternative, the processor and the storage medium may reside as discrete components in a user terminal.

[0157] In accordance with the description herein, suitable computing devices include, by way of non-limiting examples, server computers, desktop computers, laptop computers, notebook computers, sub-notebook computers, netbook computers, netpad computers, set-top computers, media streaming devices, handheld computers, Internet appliances, mobile smartphones, tablet computers, personal digital assistants, video game consoles, and vehicles. Various televisions, video players, and digital music players with optional computer network connectivity are suitable for use in the system described herein. Suitable tablet computers, in various embodiments, include those with booklet, slate, and convertible configurations.

[0158] In some embodiments, the computing device includes an operating system configured to perform executable instructions. The operating system is, for example, software, including programs and data, which manages the device’s hardware and provides services for execution of applications. Various suitable server operating systems include, by way of non-limiting examples, FreeBSD, OpenBSD, NetBSD®, Linux, Apple® Mac OS X Server®, Oracle® Solaris®, Windows Server®, and Novell® NetWare®. Various suitable personal computer operating systems include, by way of non-limiting examples, Microsoft® Windows®, Apple® Mac OS X®, UNIX®, and UNIX-like operating systems such as GNU / Linux®. In some embodiments, the operating system is provided by cloud computing. Various suitable mobile smartphone operating systems include, by way of non-limiting examples, Nokia® Symbian® OS, Apple® iOS®, Research In Motion® BlackBerry OS®, Google® Android®, Microsoft® Windows Phone® OS, Microsoft® Windows Mobile® OS, Linux®, and Palm® WebOS®. Various suitable media streaming device operating systems include, by way of non-limiting examples, Apple TV®, Roku®, Boxee®, Google TV®, Google Chromecast®, Amazon Fire®, and Samsung® HomeSync®. Various suitable video game console operating systems include, by way of nonlimiting examples, Sony® PS3®, Sony® PS4®, Microsoft® Xbox 360®, Microsoft Xbox One, Nintendo® Wii®, Nintendo® Wii U®, and Ouya®.Non-transitory computer readable storage medium

[0159] In some embodiments, the platforms, systems, media, and methods disclosed herein include one or more non-transitory computer readable storage media encoded with a program including instructions executable by the operating system of an optionally networked computing device. In further embodiments, a computer readable storage medium is a tangible component ofAttomey Docket No. 70528-701.601a computing device. In still further embodiments, a computer readable storage medium is optionally removable from a computing device. In some embodiments, a computer readable storage medium includes, by way of non-limiting examples, CD-ROMs, DVDs, flash memory devices, solid state memory, magnetic disk drives, magnetic tape drives, optical disk drives, distributed computing systems including cloud computing systems and services, and the like. In some cases, the program and instructions are permanently, substantially permanently, semipermanently, or non-transitorily encoded on the media.Computer program

[0160] In some embodiments, the platforms, systems, media, and methods disclosed herein include at least one computer program, or use of the same. A computer program includes a sequence of instructions, executable by one or more processor(s) of the computing device’s CPU, written to perform a specified task. Computer readable instructions may be implemented as program modules, such as functions, objects, Application Programming Interfaces (APIs), computing data structures, and the like, that perform particular tasks or implement particular abstract data types. A computer program may be written in various versions of various languages.

[0161] The functionality of the computer readable instructions may be combined or distributed as desired in various environments. In some embodiments, a computer program comprises one sequence of instructions. In some embodiments, a computer program comprises a plurality of sequences of instructions. In some embodiments, a computer program is provided from one location. In other embodiments, a computer program is provided from a plurality of locations. In various embodiments, a computer program includes one or more software modules. In various embodiments, a computer program includes, in part or in whole, one or more web applications, one or more mobile applications, one or more standalone applications, one or more web browser plug-ins, extensions, add-ins, or add-ons, or combinations thereof.Web application

[0162] In some embodiments, a computer program includes a web application. A web application, in various embodiments, utilizes one or more software frameworks and one or more database systems. In some embodiments, a web application is created upon a software framework such as Microsoft® .NET or Ruby on Rails (RoR). In some embodiments, a web application utilizes one or more database systems including, by way of non-limiting examples, relational, non-relational, object oriented, associative, XML, and document oriented database systems. In further embodiments, suitable relational database systems include, by way of nonlimiting examples, Microsoft® SQL Server, mySQL™, and Oracle®. A web application, inAttomey Docket No. 70528-701.601various embodiments, is written in one or more versions of one or more languages. A web application may be written in one or more markup languages, presentation definition languages, client-side scripting languages, server-side coding languages, database query languages, or combinations thereof. In some embodiments, a web application is written to some extent in a markup language such as Hypertext Markup Language (HTML), Extensible Hypertext Markup Language (XHTML), or extensible Markup Language (XML). In some embodiments, a web application is written to some extent in a presentation definition language such as Cascading Style Sheets (CSS). In some embodiments, a web application is written to some extent in a client-side scripting language such as Asynchronous JavaScript and XML (AJAX), Flash® ActionScript, JavaScript, or Silverlight®. In some embodiments, a web application is written to some extent in a server-side coding language such as Active Server Pages (ASP), ColdFusion®, Perl, Java™, JavaServer Pages (JSP), Hypertext Preprocessor (PHP), Python™, Ruby, Tel, Smalltalk, WebDNA®, or Groovy. In some embodiments, a web application is written to some extent in a database query language such as Structured Query Language (SQL). In some embodiments, a web application integrates enterprise server products such as IBM® Lotus Domino®. In some embodiments, a web application includes a media player element. In various further embodiments, a media player element utilizes one or more of many suitable multimedia technologies including, by way of non-limiting examples, Adobe® Flash®, HTML 5, Apple® QuickTime®, Microsoft® Silverlight®, Java™, and Unity®.

[0163] Referring to FIG. 2, in a particular embodiment, an application provision system comprises one or more databases 200 accessed by a relational database management system (RDBMS) 210. Suitable RDBMSs include Firebird, MySQL, PostgreSQL, SQLite, Oracle Database, Microsoft SQL Server, IBM DB2, IBM Informix, SAP Sybase, Teradata, and the like. In this embodiment, the application provision system further comprises one or more application severs 220 (such as Java servers, .NET servers, PHP servers, and the like) and one or more web servers 230 (such as Apache, IIS, GWS and the like). The web server(s) optionally expose one or more web services via app application programming interfaces (APIs) 240. Via a network, such as the Internet, the system provides browser-based and / or mobile native user interfaces.

[0164] Referring to FIG. 3, in a particular embodiment, an application provision system alternatively has a distributed, cloud-based architecture 300 and comprises elastically load balanced, auto-scaling web server resources 310 and application server resources 320 as well synchronously replicated databases 330.Mobile application

[0165] In some embodiments, a computer program includes a mobile application provided to a mobile computing device. In some embodiments, the mobile application is provided to a mobileAttomey Docket No. 70528-701.601computing device at the time it is manufactured. In other embodiments, the mobile application is provided to a mobile computing device via the computer network described herein.

[0166] In view of the disclosure provided herein, a mobile application is created by various techniques using hardware, languages, and development environments. Mobile applications may be written in various languages. Suitable programming languages include, by way of nonlimiting examples, C, C++, C#, Objective-C, Java™, JavaScript, Pascal, Object Pascal, Python™, Ruby, VB.NET, WML, and XHTML / HTML with or without CSS, or combinations thereof.

[0167] Suitable mobile application development environments are available from several sources. Commercially available development environments include, by way of non-limiting examples, AirplaySDK, alcheMo, Appcelerator®, Celsius, Bedrock, Flash Lite, .NET Compact Framework, Rhomobile, and WorkLight Mobile Platform. Other development environments are available without cost including, by way of non-limiting examples, Lazarus, MobiFlex, MoSync, and Phonegap. Also, mobile device manufacturers distribute software developer kits including, by way of non-limiting examples, iPhone and iPad (iOS) SDK, Android™ SDK, BlackBerry® SDK, BREW SDK, Palm® OS SDK, Symbian SDK, webOS SDK, and Windows® Mobile SDK.

[0168] Various commercial forums are available for distribution of mobile applications including, by way of non-limiting examples, Apple® App Store, Google® Play, Chrome WebStore, BlackBerry® App World, App Store for Palm devices, App Catalog for webOS, Windows® Marketplace for Mobile, Ovi Store for Nokia® devices, Samsung® Apps, and Nintendo® DSi Shop.Standalone application

[0169] In some embodiments, a computer program includes a standalone application, which is a program that is run as an independent computer process, not an add-on to an existing process, e.g., not a plug-in. Standalone applications may be compiled. A compiler is a computer program(s) that transforms source code written in a programming language into binary object code such as assembly language or machine code. Suitable compiled programming languages include, by way of non-limiting examples, C, C++, Objective-C, COBOL, Delphi, Eiffel, Java™, Lisp, Python™, Visual Basic, and VB .NET, or combinations thereof. Compilation is often performed, at least in part, to create an executable program. In some embodiments, a computer program includes one or more executable complied applications.Web browser plug-in

[0170] In some embodiments, the computer program includes a web browser plug-in (e.g., extension, etc.). In computing, a plug-in is one or more software components that add specificAttomey Docket No. 70528-701.601functionality to a larger software application. Makers of software applications support plug-ins to enable third-party developers to create abilities which extend an application, to support easily adding new features, and to reduce the size of an application. When supported, plug-ins enable customizing the functionality of a software application. For example, plug-ins are commonly used in web browsers to play video, generate interactivity, scan for viruses, and display particular file types. Various web browser plug-ins include Adobe® Flash® Player, Microsoft® Silverlight®, and Apple® QuickTime®. In some embodiments, the toolbar comprises one or more web browser extensions, add-ins, or add-ons. In some embodiments, the toolbar comprises one or more explorer bars, tool bands, or desk bands.

[0171] Various plug-in frameworks are suitable for development of plug-ins in various programming languages, including, by way of non-limiting examples, C++, Delphi, Java™, PHP, Python™, and VB .NET, or combinations thereof.

[0172] Web browsers (also called Internet browsers) are software applications, designed for use with network-connected computing devices, for retrieving, presenting, and traversing information resources on the World Wide Web. Suitable web browsers include, by way of nonlimiting examples, Microsoft® Internet Explorer®, Mozilla® Firefox®, Google® Chrome, Apple® Safari®, Opera Software® Opera®, and KDE Konqueror. In some embodiments, the web browser is a mobile web browser. Mobile web browsers (also called microbrowsers, mini-browsers, and wireless browsers) are designed for use on mobile computing devices including, by way of nonlimiting examples, handheld computers, tablet computers, netbook computers, subnotebook computers, smartphones, music players, personal digital assistants (PDAs), and handheld video game systems. Suitable mobile web browsers include, by way of non-limiting examples, Google® Android® browser, RIM BlackBerry® Browser, Apple® Safari®, Palm® Blazer, Palm® WebOS® Browser, Mozilla® Firefox® for mobile, Microsoft® Internet Explorer® Mobile, Amazon® Kindle® Basic Web, Nokia® Browser, Opera Software® Opera® Mobile, and Sony® PSP™ browser.Software modules

[0173] In some embodiments, the platforms, systems, media, and methods disclosed herein include software, server, and / or database modules, or use of the same. In view of the disclosure provided herein, software modules are created by various techniques using machines, software, and languages. The software modules disclosed herein are implemented in a multitude of ways. In various embodiments, a software module comprises a file, a section of code, a programming object, a programming structure, a distributed computing resource, a cloud computing resource, or combinations thereof. In further various embodiments, a software module comprises a plurality of files, a plurality of sections of code, a plurality of programming objects, a pluralityAttomey Docket No. 70528-701.601of programming structures, a plurality of distributed computing resources, a plurality of cloud computing resources, or combinations thereof. In various embodiments, the one or more software modules comprise, by way of non-limiting examples, a web application, a mobile application, a standalone application, and a distributed or cloud computing application. In some embodiments, software modules are in one computer program or application. In other embodiments, software modules are in more than one computer program or application. In some embodiments, software modules are hosted on one machine. In other embodiments, software modules are hosted on more than one machine. In further embodiments, software modules are hosted on a distributed computing platform such as a cloud computing platform. In some embodiments, software modules are hosted on one or more machines in one location. In other embodiments, software modules are hosted on one or more machines in more than one location.Databases

[0174] In some embodiments, the platforms, systems, media, and methods disclosed herein include one or more databases, or use of the same. Various databases are suitable for storage and retrieval of genetic data and / or transcriptomic data. In various embodiments, suitable databases include, by way of non-limiting examples, relational databases, non-relational databases, object oriented databases, object databases, entity-relationship model databases, associative databases, XML databases, document oriented databases, and graph databases. Further non-limiting examples include SQL, PostgreSQL, MySQL, Oracle, DB2, Sybase, and MongoDB. In some embodiments, a database is Internet-based. In further embodiments, a database is web-based. In still further embodiments, a database is cloud computing-based. In a particular embodiment, a database is a distributed database. In other embodiments, a database is based on one or more local computer storage devices.III. KITS

[0175] Disclosed herein, in some embodiments, are kits for sample collection and / or preparation that may be useful for performing the methods of the present disclosure. In some embodiments, the kit comprises reagents. In some embodiments, the reagenss are configured to preserve the sample. In some embodiments, the reagents are configured for use in at least one of a transcriptomic assay or a genotyping assay. In some embodiments, the kits comprise at least one buffer. In some embodiments, the kit comprises at least one packaging components. In some embodiments, the at least one packaging component , applicators, etc. for sample collection and preparation, and methods of detection]

[0176] The exact nature of the components configured in the inventive kit depends on its intended purpose. In some embodiments, the intended purpose for use in a healthcare setting. InAttomey Docket No. 70528-701.601some embodiments, the kit is configured to be used, at least in part, in a location outside of a health care facility. In some embodiments, the kit is configured to be used, at least in part, by the subject.

[0177] Instructions for use may be included in the kit. “Instructions for use” typically include a tangible expression describing the technique to be employed in using the components of the kit to effect a desired outcome, such as to determine an output indicative of the subject as having a degree of cancer risk or to treat the cancer. Optionally, the kit also contains other useful components, such as, diluents, buffers, pharmaceutically acceptable carriers, syringes, catheters, applicators, pipetting or measuring tools, bandaging materials or other useful paraphernalia as will be readily recognized by those of skill in the art.

[0178] The materials or components assembled in the kit can be provided to the practitioner stored in any convenient and suitable ways that preserve their operability and utility. For example, the components can be in dissolved, dehydrated, or lyophilized form; they can be provided at room, refrigerated or frozen temperatures. The components are typically contained in suitable packaging material(s). As employed herein, the phrase “packaging material” refers to one or more physical structures used to house the contents of the kit, such as inventive compositions and the like. The packaging material is constructed by well-known methods, preferably to provide a sterile, contaminant-free environment. The packaging materials employed in the kit are those customarily utilized in gene expression assays and in the administration of treatments. As used herein, the term “package” refers to a suitable solid matrix or material such as glass, plastic, paper, foil, and the like, capable of holding the individual kit components. The packaging material generally has an external label which indicates the contents and / or purpose of the kit and / or its components.

[0179] Disclosed herein, are kits useful for to detect the genotypes and / or biomarkers disclosed herein. In some embodiments, the kits disclosed herein may be used to diagnose and / or treat a disease or condition in a subject; or select a patient for treatment and / or monitor a treatment disclosed herein. In some embodiments, the kit comprises the compositions described herein, which can be used to perform the methods described herein. In other embodiments, the kits contains all of the components necessary and / or sufficient to perform an assay for detecting and measuring bladder cancer markers, including all controls, directions for performing assays, and any necessary software for analysis and presentation of results.

[0180] In some instances, the kits described herein comprise components for detecting the presence, absence, and / or quantity of a target nucleic acid and / or protein described herein. In some embodiments, the kit further comprises components for detecting the presence, absence, and / or quantity of a serological marker described herein. In some embodiments, the kitAttomey Docket No. 70528-701.601comprises the compositions (e.g., primers, probes, antibodies) described herein. The disclosure provides kits suitable for assays such as enzyme-linked immunosorbent assay (ELISA), single- molecular array (Simoa), PCR, qPCR, quantitative rtPCR, and ddPCR. The exact nature of the components configured in the kit depends on its intended purpose.

[0181] In some embodiments, the kits described herein are configured for the purpose of treating and / or characterizing a disease or condition (e.g., bladder cancer or UC), or subclinical phenotype thereof (e.g., muscle invasive or non-muscle invasive) in a subject. In some embodiments, the kits described herein are configured for the purpose of to determining an output indicative of the subject as having a degree of cancer risk a subject. In some embodiments, the kit is configured particularly for the purpose of treating mammalian subjects. In some embodiments, the kit is configured particularly for the purpose of treating human subjects. In further embodiments, the kit is configured for veterinary applications, treating subjects such as, but not limited to, farm animals, domestic animals, and laboratory animals.

[0182] Instructions for use may be included in the kit. Optionally, the kit also contains other useful components, such as, diluents, buffers, pharmaceutically acceptable carriers, syringes, catheters, applicators, pipetting or measuring tools, bandaging materials or other useful paraphernalia. The materials or components assembled in the kit can be provided to the practitioner stored in any convenient and suitable ways that preserve their operability and utility. For example the components can be in dissolved, dehydrated, or lyophilized form; they can be provided at room, refrigerated or frozen temperatures. The components are typically contained in suitable packaging material(s). As employed herein, the phrase “packaging material” refers to one or more physical structures used to house the contents of the kit, such as compositions and the like. The packaging material is constructed by well-known methods, preferably to provide a sterile, contaminant-free environment. The packaging materials employed in the kit are those customarily utilized in gene expression assays and in the administration of treatments. As used herein, the term “package” refers to a suitable solid matrix or material such as glass, plastic, paper, foil, and the like, capable of holding the individual kit components. Thus, for example, a package can be a glass vial or prefilled syringes used to contain suitable quantities of the pharmaceutical composition. The packaging material has an external label which indicates the contents and / or purpose of the kit and its components.IV. DEFINITIONS

[0183] Unless defined otherwise, all terms of art, notations and other technical and scientific terms or terminology used herein are intended to have the same meaning as is commonly understood by one of ordinary skill in the art to which the claimed subject matter pertains. InAttomey Docket No. 70528-701.601some cases, terms with commonly understood meanings are defined herein for clarity and / or for ready reference, and the inclusion of such definitions herein should not necessarily be construed to represent a substantial difference over what is generally understood in the art.

[0184] Throughout this application, various embodiments may be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the disclosure.Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, for example, 1, 2, 3, 4, 5, and 6. This applies regardless of the breadth of the range.

[0185] As used in the specification and claims, the singular forms “a”, “an” and “the” include plural references unless the context clearly dictates otherwise. For example, the term “a sample” includes a plurality of samples, including mixtures thereof.

[0186] The terms “determining,” “measuring,” “evaluating,” “assessing,” “assaying,” and “analyzing” are often used interchangeably herein to refer to forms of measurement. The terms include determining if an element is present or not (for example, detection). These terms can include quantitative, qualitative or quantitative and qualitative determinations. Assessing can be relative or absolute. “Detecting the presence of’ can include determining the amount of something present in addition to determining whether it is present or absent depending on the context.

[0187] The terms “subject,” “individual,” or “patient” are often used interchangeably herein. A “subject” can be a biological entity containing expressed genetic materials. The biological entity can be a plant, animal, or microorganism, including, for example, bacteria, viruses, fungi, and protozoa. The subject can be tissues, cells and their progeny of a biological entity obtained in vivo or cultured in vitro. The subject can be a mammal. The mammal can be a human. The subject may be diagnosed or suspected of being at high risk for a disease. In some cases, the subject is not necessarily diagnosed or suspected of being at high risk for the disease.

[0188] The term “zzz vivo" is used to describe an event that takes place in a subject’s body.

[0189] The term “ex vivo" is used to describe an event that takes place outside of a subject’s body. An ex vivo assay is not performed on a subject. Rather, it is performed upon a sample separate from a subject. An example of an ex vivo assay performed on a sample is an “zzz vitro" assay.Attomey Docket No. 70528-701.601

[0190] The term “ / / / vitro" is used to describe an event that takes places contained in a container for holding laboratory reagent such that it is separated from the biological source from which the material is obtained. In vitro assays can encompass cell-based assays in which living or dead cells are employed. In vitro assays can also encompass a cell-free assay in which no intact cells are employed.

[0191] As used herein, the term “about” a number refers to that number plus or minus 10% of that number. The term “about” a range refers to that range minus 10% of its lowest value and plus 10% of its greatest value.

[0192] As used herein, the terms “treatment” or “treating” are used in reference to a pharmaceutical or other intervention regimen for obtaining beneficial or desired results in the recipient. Beneficial or desired results include but are not limited to a therapeutic benefit and / or a prophylactic benefit. A therapeutic benefit may refer to eradication or amelioration of symptoms or of an underlying disorder being treated. Also, a therapeutic benefit can be achieved with the eradication or amelioration of one or more of the physiological symptoms associated with the underlying disorder such that an improvement is observed in the subject, notwithstanding that the subject may still be afflicted with the underlying disorder. A prophylactic effect includes delaying, preventing, or eliminating the appearance of a disease or condition, delaying or eliminating the onset of symptoms of a disease or condition, slowing, halting, or reversing the progression of a disease or condition, or any combination thereof. For prophylactic benefit, a subject at risk of developing a particular disease, or to a subject reporting one or more of the physiological symptoms of a disease may undergo treatment, even though a diagnosis of this disease may not have been made.

[0193] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.V. EXAMPLES

[0194] The following examples are included for illustrative purposes only and are not intended to limit the scope of the inventive concepts.Example 1: Overview of the Identification of Biomarkers

[0195] There are five messenger RNA (mRNA) markers associated with bladder cancer, including urothelial carcinoma (UC). Four of the five mRNA markers are associated with UC: Cyclin-dependent kinase 1 (CDK1), midkine (MDK), insulin-like growth factor binding protein 5 (IGFBP5), and Homeobox Al 3 (H0XA13); and one mRNA marker is associated with inflammation: C-X-C motif chemokine receptor 2 (CXCR2). Additionally, six DNA singlenucleotide polymorphisms (SNPs) from fibroblast growth factor receptor 3 (FGFR3) andAttorney Docket No. 70528-701.601telomerase reverse transcriptase (TERT), are associated with bladder cancer, including UC. Four SNPs are from with FGFR3 (rsl21913482, rsl21913483, rsl21913479, rsl21913485), and two SNPs are from TERT (rsl242535815, rsl561215364). Table 1A illustrates the PCR Primers and probes that were used for detection of genetic markers and Table IB shows the primers and probes for FGRR Forward and reverse, and the primers and probes for detecting the TERT SNPs. Tables 1C-1H show varies conditions for anlayis including: Reagents and Concentrations for Digital PCR Assays (Table 1C), Reagents Used for TERT Emulsion ddPCR (Table ID), Thermocycling conditions(Table IE), Emulsion dPCR Reagents for FGFR3 Analysis (Table IF), Reagents for Microfluidic Array PCR for FGFR3 (Table 1G), and Thermocycling Conditions for Alternative FGFR3 Analysis (Table 1H).Table 1A. PCR Primers and Probes Used for Detection of Genetic Markers for Urothelial CancerTable IB. FGFR3 Forward (FW) and Reverse (RW) Pimers and Probes & Primers and Probes for Detecting TERT SNPsAttorney Docket No. 70528-701.601Table 1C. Reagents and Concentrations for Digital PCR AssaysAttorney Docket No. 70528-701.601Table ID. Reagents Used for TERT Emulsion ddPCRTable IE. Thermocycling conditionsTable IF. Emulsion dPCR Reagents for FGFR3 AnalysisTable 1G. Reagents for Microfluidic Array PCR for FGFR3Attorney Docket No. 70528-701.601Table 1H. Thermocycling Conditions for Alternative FGFR3 AnalysisExample 2. Training a Machine Learning Algorithm for Multi-Modal Diagnostic / Biomarker Test

[0196] The trained machine learning algorithm comprises a trained machine learning classifier, wherein the degree of cancer risk comprises a categorical cancer risk selected from among a plurality of distinct categorical cancer risks, and wherein determining the output comprises assigning a classification of the subject as having the categorical cancer risk selected from among the plurality of distinct categorical cancer risks.

[0197] To train the trained machine learning algorithm, first genotype data and transcriptomic data from training samples comprising case samples obtained from subjects with the bladder cancer were received, as well as control samples obtained from subjects without the bladder cancer were received. Then a first training set was created, wherein the first training set includes covariate data corresponding to the genotype data and the transcriptomic data. Then, the machine learning algorithm was trained with the covariate data to determine an output indicative of the subject as having a degree of cancer risk.Example 3: Validation of the Machine Learning Algorithm for Multi-Modal Diagnostic / Biomarker Test

[0198] The trained machine learning workflow identified an algorithm for use in the multimodal diagnostic assay in Example 2. In order to validate the findings, a cohort of patients with hematuria were tested using the assay, and the results compared to cystoscopy results for aAttomey Docket No. 70528-701.601confirmed positive or negative diagnosis of urothelial cancer (UC). Patients with confirmed hematuria were recruited. Eligible patients were being evaluated by flexible or rigid cystoscopy as clinically indicated, including those who were referred due to suspicious or positive imaging, and were able to provide a voided urine sample (>30 mL) and comply with study requirements. Midstream urine samples were collected from patients prior to cystoscopy or >14 days after their prior diagnostic cystoscopy procedure. The samples were collected and processed using the Cxbladder Urine Sampling System and analyzed by Pacific Edge Diagnostics New Zealand (PEDNZ).

[0199] An analysis cohort was recruited. A total of 702 patients with hematuria provided samples. Of the 702 patients, 87 were excluded from the analysis cohort for reasons including a prior history of bladder malignancy, prostate or renal cell carcinoma, prior genitourinary manipulation within 14 days of urine sample collection, alkylating agent-based chemotherapy, current pregnancy, or other protocol deviations. Other reasons for exclusion from the analysis cohort included: a rejection of the urine sample due to not meeting PEDNZ laboratory quality control standards, a screen failure, unconfirmed tumors (i.e., no pathological confirmation of tumor), patient consent was verified too late, or no cystoscopy was performed within 90 days. Of the 702 patients, 615 patients were included in the analysis cohort.

[0200] In order to move forward with analyzing the diagnostic test, the analysis set was compared to the cystoscopy results for a confirmed positive or negative diagnosis of UC. A confirmed positive or negative diagnosis of UC was determined according to the following criteria: patients with a pathological confirmation of UC (from biopsied or resected tissue from cystoscopy and / or ureteroscopy) were considered UC positive; patients who had no abnormalities on cystoscopy or identified from biopsied or resected tissue from cystoscopy and / or ureteroscopy were considered UC negative; patients who underwent treatment that ablated tissue without biopsy were classified as UC positive without confirmation (such patients were documented fully and excluded from analysis); and patients who were diagnosed with papillary urothelial neoplasm of low malignant potential (PUNLMP) were considered UC negative for this study, with PUNLMP considered as an alternative diagnosis for the hematuria event. Pathology-confirmed tumors were classified as high grade (HG; i.e., high grade, stage >Ta) or carcinoma in situ (Cis) tumors, or low grade (LG; i.e., low grade, Ta, unknown). AUA risk stratification was conducted retrospectively based on available data.

[0201] Next, the analysis cohort was tested by the diagnostic test to compute a score. Of the 615 patients of the analysis cohort, 28 did not yield a result from the diagnostic test. Reasons for no result include inflammation, technical replicate failure, insufficient DNA, or qPCR failure. FIG.6 shows a flow diagram of patient selection for the analysis set. In this example, the algorithmAttorney Docket No. 70528-701.601used two pre-specified score thresholds to indicate the results as either low, intermediate, or a high risk result. For scores lower than 0.15, the risk was determined to be low; for score between 0.15-0.54, the risk was determined to be intermediate; for scores greater than 0.54, the risk was high. A calibration curve representing these pre-specified thresholds are shown in FIG. 7. For scores that are low-risk, the result of the diagnostic test was negative. For scores that are intermediate- or high-risk, the result of the diagnostic test was positive. Of the 587 patients with results from the multi-modal diagnostic workflow, 171 patients (29.1%) tested Intermediate or High risk of UC and 416 (70.9%) tested Low risk of UC (i.e., ‘ruled out’), as seen in Table 2.Table 2: Confusion Matrix of the Multi-Modal Diagnostic Workflow Versus Confirmed Tumor Status

[0202] Table 2 also shows that of the ‘ruled out’ patients, 413 / 416 (99.3%) had normal cystoscopy; three tumors (one HG and two LG) were missed.Example 4: Calculation of Performance Characteristics of the Multi-Modal Diagnostic Workflow

[0203] The multi-modal diagnostic workflow was validated using the results from Example 3. Calculations of sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), and a test-negative rate (TNR) were calculated from the 587 diagnostic test results compared with cystoscopy as described above. As shown in Table 3, the multi-modal diagnostic workflow had sensitivity of 94% (95% confidence interval 83-99%), specificity of 77% (73-80%), positive predictive value (PPV) of 26% (20-34%), negative predictive value (NPV) of 99.3% (99.3-99.9%), and a test-negative rate (TNR) of 71% (67-75%). These performances were similar in patients with gross hematuria (GH) or microhematuria (MH).Attorney Docket No. 70528-701.601Table 3. Performance Calculations of Multi-Modal Diagnostic WorkflowExample 5: The Multi-Modal Diagnostic Workflow Compared with AUA & Hematuria Status Risk Stratification

[0204] AUA risk stratification was conducted retrospectively on the analytic set from Example 1, based on available data. Similarly, hematuria-based risk stratification was also conducted on the analytic set from Example 1. Table 4 shows that 7 / 584 (1.02%) of patients were given a low risk by following AUA guidelines, 88 / 584 (15.1%) of patients had an AUA intermediate risk, and 489 / 584 (83.7%) of patients had a high AUA risk. Pathology-confirmed UC tumors were detected in 48 / 489 (9.8%) AUA high-risk patients, compared with 45 / 171 (26.3%) multi-modal diagnostic workflow-positive cases. This corresponds to a 2.7-fold (26.3% / 9.8%) improvement in PPV with the multi-modal diagnostic workflow over AUA risk stratification.Attorney Docket No. 70528-701.601Table 4: Confusion Matrix of the AUA Risk Stratification Results Versus Confirmed Tumor Status

[0205] Table 5 shows that 267 / 586 patients had gross hematuria and 320 patients have microhematuria. When using hematuria status to indicate risk, confirmed tumors were detected in 30 / 267 (11.2%) patients with GH, as seen in Table 5. Therefore, the multi-modal diagnostic workflow showed a 2.3-fold (26.3% / l 1.2%) improvement in PPV compared with hematuriabased risk stratification.Table 5: Confusion Matrix of the Hematuria Risk Stratification Results Versus Confirmed Tumor StatusExample 5: Comparison of the Multi-Modal Diagnostic Workflow with A Frontline Diagnostic & the Secondary Diagnostic Assays

[0206] The analytic set from Example 1 was tested using two additional tests: a frontline diagnostic and the secondary diagnostic. Performance characteristics of each test were then calculated,sss similar to the performance characteristics of Example 3. The multi-modal diagnostic workflow and performance characteristics from Example 3 were then compared against the performance characteristics of the frontline diagnostic and the secondary diagnostic. A frontline diagnosticAttorney Docket No. 70528-701.601

[0207] First, the analysis cohort was tested by A frontline diagnostic, which outputs a Negative or Positive result. Of the 615 samples tested by A frontline diagnostic, 205 had a negative result, 375 had a positive result, 32 samples had no result, and 3 results were not available due to missing clinical data. Like in Example 2, the frontline diagnostic results were compared to the cystoscopy results for a confirmed positive or negative diagnosis of UC (i.e., tumor status). Then, the comparison of the frontline diagnostic to tumor status was further compared to the Multi-Modal Diagnostic Workflow results in Table 6.Table 6. Confusion Matrix of the Multi-Modal Diagnostic Workflow and A frontline diagnostic Versus Confirmed Tumor Status

[0208] Next, calculations of performance characteristics were calculated for the frontline diagnostic, similar to Example 3. Calculations of sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), and a test-negative rate (TNR) were calculated from the 580 diagnostic test results compared with cystoscopy as described above. As shown in Table 7, the frontline diagnostic had sensitivity of 93% (95% confidence interval 82-99%), specificity of 38% (34-42%), positive predictive value (PPV) of 11% (8-15%), negative predictive value (NPV) of 98.5% (95.8-99.7%), and a test-negative rate (TNR) of 35% (31-39%). The performance characteristics of the frontline diagnostic were then compared to the multi-modal diagnostic workflow performance characteristics from Example 3, with the results also being shown in Table 7. When the frontline diagnostic was used, 375 / 615 (61.0%) patients tested High risk, and 42 / 48 (87.5%) tumors were detected. Compared with the multi-modalAttorney Docket No. 70528-701.601diagnostic workflow, The frontline diagnostic had similar sensitivity and NPV, but lower specificity, PPV, and TNR. Additionally, Receiver Operating characteristic (ROC) curves of the frontline diagnostic and the multi-modal diagnostic workflow were plotted to compare the performance of the two tests, as shown in FIG. 8. The ROC curves showed a higher sensitivity and a lower false-positive rate with the multi-modal diagnostic workflow compared to the frontline diagnostic.Table 7: Performance Characteristics of the Multi-Modal Diagnostic Workflow Compared with Performance Characteristics of The frontline diagnosticAttorney Docket No. 70528-701.601The secondary diagnostic

[0209] The comparison of the secondary diagnostic with the multi-modal detection workflow follows similar steps as the comparison with the frontline diagnostic. First, the analysis cohort was tested by Detect, which outputs a Negative or Positive result. Of the 615 samples tested by Detect, 416 had a negative result, 122 had a positive result, and 29 samples had no result. Like in Example 2, the secondary diagnostic results were compared to the cystoscopy results for a confirmed positive or negative diagnosis of UC (i.e., tumor status). Then, the comparison of the secondary diagnostic to tumor status was further compared to the Multi-Modal Diagnostic Workflow results in Table 8.Table 8. Confusion Matrix of the Multi-Modal Diagnostic Workflow and The Secondary Diagnostic Versus Confirmed Tumor Status

[0210] Next, calculations of performance characteristics were calculated for Detect, similar to Example 3. Calculations of sensitivity, specificity, positive predictive value (PPV), negativeAttorney Docket No. 70528-701.601predictive value (NPV), and a test-negative rate (TNR) were calculated from the 580 diagnostic test results compared with cystoscopy as described above. As shown in Table 9, The secondary diagnostichad sensitivity of 65% (95% confidence interval 50-79%), specificity of 77% (74-81%), positive predictive value (PPV) of 20% (14-27%), negative predictive value (NPV) of 96.3% (94.1-97.9%), and a test-negative rate (TNR) of 74% (70-77%). The performance characteristics of the secondary diagnosticwere then compared to the multi-modal diagnostic workflow performance characteristics from Example 3, with the results also being shown in Table 9. When the secondary diagnosticwas used, 152 / 615 (24.7%) patients tested High risk, and 30 / 48 (62.5%) tumors were detected. Compared with the multi-modal diagnostic workflow, The secondary diagnostichad similar specificity, PPV, and TNR, but lower sensitivity and NPV.Table 9: Performance Characteristics of the Multi-Modal Diagnostic Workflow Compared with Performance Characteristics of DetectAttorney Docket No. 70528-701.601Example 6: Analytical Validation of Six DNA SNPs of the Multi-Modal Diagnostic Workflow

[0211] A developmental dataset was gathered to conduct the analytical validation of the six DNA single-nucleotide polymorphisms (SNPs) from FGFR3 (rsl21913482, rsl21913483, rsl21913479, rsl21913485) and 7EAT(rsl242535815, rsl561215364). The dataset included stabilized urine samples stored at -80C from patients with gross hematuria or microhematuria who had participated in a previous clinical validation study and patients with hematuria who participated in the STRATA: Safe Testing of Risk for Asymptomatic Microhematuria trial. The samples were blinded prior to their use in the development dataset. A total of 987 samples were included in the dataset.Linearity

[0212] The linearity of the multi-modal diagnostic workflow of Example 1 was assessed for each of the six DNA SNPs o FGFR3 (rsl21913482, rsl21913483, rsl21913479, rsl21913485) and TERT (rs!242535815, rs!561215364), with an upper limit being 2.21 L and 1.76 , respectively. To confirm the linearity of each analytic target, a 10-point dilution curve (33,000, 6600, 1320, 264, 132, 66, 33, 16.5, 8.25, and 4.12 DNA copies / well) was performed, with at least four replicates at each concentration. Then, statistical testing was performed to determine the concentration at which the assay became non-linear for each analytic target. Data were fitted using a linear regression model, and linearity was assessed using the R2and mean squared error (MSE). The null hypothesis of this example was that the multi-modal diagnostic workflow had valid linear regression through the whole diluted range; this null hypothesis was rejected if R2was <0.9. Then, the MSE was used to compare linearity between different analytic targets.

[0213] The multi-modal diagnostic workflow of Example 1 demonstrated linearity across all six FGFR3 and TERT SNPs with an R2value of >0.99 for all analytic targets (as seen in FIG. 5).The MSE values for individual SNPs ranged from 0.005 to 0.014. The tested range was 1 x 105Attomey Docket No. 70528-701.601to 1x101DNA copies / well, which corresponded to a maximum linear concentration of 5k. The results of this example demonstrate that the linearity of this assay mt the criteria for linearity in diagnostic testing.Analytical Sensitivity

[0214] Next, the analytical sensitivity, or the limit of detection (LOD), the multi-modal diagnostic workflow was evaluated. Analytical sensitive was defined as the lowest analyte concentration that could be consistently detected with 95% probability. The LOD for the DNA SNPs of FGFR3 and TERT was determined by the logistic regression and the three-concentrations approach.

[0215] For logistic regression, a model was fitted to explain the relationship between the dependent (positive / negative) and independent (analyte concentration) variables, and the LOD was predicted based on this fitted logistic model. For the three-concentrations approach, a fraction of positives was first calculated from highest to lowest concentration, and the concentration that matched 95% of positives (cO) was the boundary of LOD.

[0216] Samples with concentrations >c0 were grouped into three concentration levels: cl, c2, and c3, where cl > c2 > c3 > cO. If ni was the number of positive samples at concentration ci, it was assumed that ni ~ Poi(kci) (where i = 1, 2, 3); the expected value E(ni) = kci was the average number of positive samples at concentration ci. If pi was the probability of no samples detecting positive at concentration ci, then pi = P(ni = 0) = e-kci. If c* was the LOD, by definition, 1— pi = 0.95 and pi = e-kc* = 0.05; hence, c* = log 0.05 / -Ak. If Xi is the observed number of positive samples at concentration ci, the maximum likelihood estimatorAk for X can be computed using the multi-modal diagnostic workflow algorithm of Example 1. The three concentrations (cl, c2, and c3) were selected in multiple ways and the corresponding number of samples at each concentration changed accordingly, with both (the number of samples and concentration) having an impact on the LOD estimate. The most conservative value was used as the estimated LOD.

[0217] The predicted LOD of the multi-modal diagnostic workflow of Example 1 is shown below in Table 10. When using the logistic approach, the predicted LOD of the multi-modal diagnostic workflow for DNA detection was 840, 1200, 1250, and 970 copies / mL for FGFR3 c.742, c.746, c.1114, and c.1124, respectively, and 440 and 740 copies / mL for TERT mtl24 and mtl46, respectively. Using the three concentrations approach, the predicted LOD was 632, 1220, 946, and 439 copies / mL for FGFR3 c.742, c.746, c.1114, and c.1124, respectively, and 319 and 418 copies / mL for TERT mtl24 and mtl46, respectively. There was no significant difference between the logistic regression and three concentrations approach for LOD prediction; however,Attorney Docket No. 70528-701.601the logistic regression approach was preferred for all SNPs except FGFR3 c.1114. The lower 95% CI for FGFR3 c.1124 was not estimable due to large variations in the data.Table. 10. Limit of Detections of 6 DNA SNPsAnalytical Specificity

[0218] Next, the analytical specificity of the multi-modal diagnostic workflow was evaluated. Analytical specificity was defined as the ability of the multi-modal diagnostic workflow to detect FGFR3 and TERT mutant DNA in the presence of potentially interfering substances, which may have been carried over from the extraction reagent or present in the patient’s urine sample. The effect on analytical specificity of the assay for both sample- and process-derived interfering substances was assessed. The sample-derived substances assessed were red blood cells (RBCs; 8 * 105, 4 x io6, 2 * 107, and 1 x io8cells / mL), bacteria (1xio6cells / mL Escherichia coli), yeast (1 x 104colony-forming units [cfu] / mL), urea (60 mg / mL), glucose (0.5 mg / mL), and protein (1.25, 2.50, 5.00, and 10.00 mg / mL serum albumin). The selected amount of each substance was mixed with eight high- and low-concentration FGFR3 and TERT extraction controls. The contaminated controls were then extracted and compared with high- and low-concentration controls without interfering substances. Samples within the expected level of gene variance and where the multi-modal diagnostic workflow score was within the 95% confidence interval (CI) for the control were considered acceptable. The difference was calculated by subtracting the mean score for the control sample from that of each contaminated sample. The corresponding p-value was calculated using a two-sided t test.Attorney Docket No. 70528-701.601

[0219] The results of the experiment for sample-derived substances are shown below in Table 11. For sample-derived substances, the results show that RBCs had no significant effect on extraction of HECs or LECs of FGFR3 and TERT at RBC concentrations of up to 2 / | 07cells / mL, with only a potential downward trend in detection in LECs. At RBC levels of 1 x 108cells / mL, there was significant impact on extraction efficiency; however, there was no sensitivity loss as all HECs returned a positive the multi-modal diagnostic workflow result. LECs had some loss of sensitivity at RBCs of 1 x 108cells / mL, with the potential to return a false-negative test result. The results show that the presence of clinically high levels of bacteria (E.coli 1 x io6cells / mL), glucose (0.5 mg / mL), urea (60 mg / mL), or yeast (1 x 104cfu / mL) had no significant effect on the extraction of HECs and LECs of FGFR3 and TERT or assay performance. However, the presence of protein had an effect on the extraction efficiency and DNA detection, with the assay being inhibited by protein levels of >2.5 mg serum albumin / mL.Table 11. Analytical specificity of the Multi-Modal Diagnostic Workflow for FGFR3 and TERT DNA control samples when mixed with sample-derived interfering substancesAttorney Docket No. 70528-701.601

[0220] The process-derived substances assessed were ethanol (4%), MagMAX wash buffer (2%), Cxbladder stabilizing reagent (1%), MagMAX magnetic beads (5%), and acetone (10%). The substance percentage was per 64 pL of elution volume and each substance was mixed with high- and low-concentrations of FGFR3 and TERT extraction controls as 12 replicates and compared with controls (without potentially interfering substances) to determine if the samples were affected by the process-derived contaminants.

[0221] For process-derived substances, the results show that all reagents showed a slight increase in assay performance, but were within the expected level of variation. The substances that most affected the assay in this example, MagMAX wash buffer and Cxbladder stabilizing reagent, are the first and second reagent in the multistep extraction and purification; therefore, the risk of these reagents being present at the elution step at concentrations that may affect the assay is very low. The substances that are closest to the elution step (acetone and MagMAX magnetic beads) did not significantly impact the assay at the high concentrations tested.Analytical Accuracy

[0222] Next, the analytical accuracy of the multi-modal diagnostic workflow of Example 1 was evaluated. Droplet digital polymerase chain reaction (ddPCR) was considered the most accurate method to define absolute quantitation of DNA in the multiplex the multi-modal diagnosticAttorney Docket No. 70528-701.601workflow assay, since there was no analytical standard for FGFR3 and TERT. To validate the analytical accuracy of the multi-modal diagnostic workflow, ddPCR of the SNPs of FGFR3 (rsl21913482, rsl21913483, rsl21913479, rsl21913485) and 7EZ?7(rsl242535815, rsl561215364) as single analytes were compared with combined analyte samples (of mutant plus wild type [WT] DNA). A combination of TERT rsl242535815 and rsl561215364 mutant DNA was also assessed.

[0223] Then, mutant DNA — alone and combined with WT DNA — were manufactured in parallel to contain equivalent concentrations of the target DNA. High-extraction controls (HECs) had a DNA concentration of ~1 x 106copies / pL and low-extraction controls (LECs) had a DNA concentration of ~1 x 104copies / pL. Each mutant was combined with WT DNA at a high (1:10) and low (1 :200) mutant:WT ratio for the HECs and LECs, respectively. Mutant DNA samples were compared with mutant + WT DNA samples for each SNP of FGFR3 and TERT in the multiplex ddPCR assays.

[0224] Accuracy was determined as the percentage of the expected mutant DNA concentration compared with the mutant + WT DNA sample. Quantification of mutant DNA within the combined sample was then reviewed against its corresponding 95% CI.

[0225] The results of the experiment are shown below in Table 12. The absolute quantification of FGFR3 mutant DNA versus mutant + WT DNA were within the 95% CI for the expected intra-plate variance for high-extraction controls of R248C and S249F / C and low-extraction controls of S249F / C and S371C. However, higher intra-plate variance was observed for the LEC of R248C, HEC of S371C, and the HEC and LEC of Y373. TERT mutant DNA versus mutant + WT DNA quantification showed acceptable intra-plate variance for HECs of C228T and C228T + C250T, but higher intra-plate variance for all LECs and the HEC of C250T.Table 12. Results of Analytical Accuracy for the Six SNPsAttorney Docket No. 70528-701.601Analytical Precision

[0226] Next, analytical precision of the multi-modal diagnostic workflow assay was evaluated. Analytical precision was defined as reproducibility within a single run or between separate runs for replicate samples. Variance in the ddPCR assay was assessed using HECs and LECs for each DNA FGFR3 and TERT mutant. HECs had a DNA concentration of ~1 x 106copies / pL and a mutant:WT ratio of 1 : 10, and LECs had a DNA concentration of ~1 x 104copies / pL and a mutant:WT ratio of 1:200.

[0227] A fitted linear random effects model was used to estimate intra-assay and inter-assay variance for the mutant fraction, presented as standard deviation (SD) and 95% confidence interval (CI); the coefficient of variation (CV%) was also calculated. Inter-assay, intra-assay, and total variance for the mutant fraction were assessed by four operators over >60 days by reviewing either 46 ddPCR plate controls split over 22 plates per FGFR3 SNP or 48 ddPCR plate controls over 23 plates per TERT SNP.

[0228] The lot-to-lot reagent variation was also assessed by testing three independent manufactures of MagMAX magnetic beads (BD2206339, BD2206338, and BD2302342) and MagMAX wash buffer (WB2405057, WB2405056, and WB2312054) for both FGFR3 and TERT using high- and low-extraction controls. Each manufacture lot was run in parallel on a single plate with >8 replicates of each control for each lot of reagents.

[0229] The results for the analytical precision of the multi-modal diagnostic workflow are shown in Table 13 below. Inter-assay variability showed mutant fraction CV% of <4.33% for HECs, with mutant fraction CV%s of 1.35-4.33% for FGFR3 and 1.01-2.11% for TERT,Attorney Docket No. 70528-701.601however, mutant fraction variance was slightly higher for LECs (mutant fraction CV%s of 4.62-6.97% and 3.42-6.86%, respectively). Based on a maximum mutant fraction CV% of 6.97% for FGFR3 C228T and 6.86% for TERT C228T + C250T, the inter-assay variability was deemed acceptable. For intra-assay variability, HECs had mutant fraction CV%s of <3.54% for FGFR3 and <2.95% for TERT . LECs had higher mutant fraction CV%s than HECs (11.56-15.90% for FGFR3 and 11.79-20.27% for TERT). Based on a maximum mutant fraction CV% of 15.90% for FGFR3 S371C and 20.27% for TERT C228T + C250T, the intra-assay variability was deemed acceptable. Total assay mutant fraction CV% for FGFR3 and TERT was <5.59% for HECs, and 12.45-17.24% and 12.27-21.40%, respectively, for LECs. The maximum total mutant fraction variance was considered excellent for high-extraction controls and acceptable for low-extraction controls.Table 13. Analytical Precision of the Multi-Modal Diagnostic WorkflowAttorney Docket No. 70528-701.601

[0230] As shown by Table 14, Lot-to-lot reagent variance was low for FGFR3 HECs (mutant fraction CV%s 1.14-2.53%), and within an acceptable range (CV%s 13.10-17.64%) for LECs. Similarly with TERT, lot-to-lot variance was low (mutant fractions CV%s 0.84-1.74%) for HECs, but was slightly higher for LECs (CV%s 11.46-32.15%).Table 14. Lot-to-Lot Reagent VarianceExtraction Efficiency

[0231] Next, the extraction efficiency was evaluated to confirm that the multi-modal diagnostic workflow of Example 1 was not biased for extraction of either FGFR3 or TERT mutant DNA,Attorney Docket No. 70528-701.601synthetic urine samples were prepared for each SNP at high (DNA concentration of ~1 x 106 copies / pL and mutant:WT ratio of 1:10) and low (DNA concentration of ~1 x 104 copies / pL and mutant:WT ratio of 1:200) mutant:WT ratios. The extraction efficiency was tested by processing eight replicates of each synthetic urine sample from extraction to ddPCR. These results were then compared with the expected mutant fraction and copy number. The absolute quantity of the extracted sample was used to define the extraction efficiency of DNA.

[0232] The results of the evaluation are shown in Table 15. The extraction efficiency of the multi-modal diagnostic workflow for FGFR3 and TERT SNPs was slightly lower in HECs versus LECs. The FGFR3 samples had a mean extraction efficiency of 72.9% for high mutant:WT ratio controls and 88.0% for low mutant:WT ratio controls. The TERT samples had a mean extraction efficiency of 83.5% for high mutant: WT ratio controls and 95.5% for low mutant:WT ratio controls. No difference was observed between extracted samples and input controls, confirming that the extraction method did not introduce sampling bias.Table 15. Extraction Efficiency of the Multi-Modal Diagnostic WorkflowAttorney Docket No. 70528-701.601Inter-Laboratory Evaluation

[0233] An inter-laboratory comparison of the multi-modal diagnostic workflow of Example 1 was conducted between the NZ laboratory (PEDNZ) and the US laboratory (PEDUSA). A random sample set from the PEDNZ validation was used to confirm the reproducibility of the multi-modal diagnostic workflow at PEDUSA. Acceptable variability was defined as achieving >80% concordance for all clinical results. The inter-laboratory comparison between PEDNZ and PEDUSA was based on data from 33 samples. The results showed there was 87.9% concordance in clinical results for the multi-modal diagnostic workflow between the two laboratories.Example 7: Calculation of Performance Characteristics of the Multi-Modal Diagnostic Workflow with a Pre-Specified Threshold of 0.54 (s3 &s4)

[0234] The multi-modal diagnostic workflow was tested using a different pre-specified threshold from Examples 1-6. The analytic set from Example 1 was tested using a threshold of 0.54. As shown in Table 16, Under the 0.54 threshold for the multi-modal diagnostic workflow, 57 / 615 patients (9.3%) had a high risk result, and 29 / 48 tumors (60.4%) were detected.Table 16. Confusion Matrix of the Multi-Modal Diagnostic Workflow at the 0.54 Threshold Versus Confirmed Tumor Status

[0235] Next, calculations of performance characteristics were calculated for the multi-modal diagnostic workflow at the 0.54 threshold. Calculations of sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), and a test-negative rate (TNR) were calculated from the 587 diagnostic test results compared with cystoscopy as described above. As shown in Table 17, had a sensitivity of 60% (95% CI 45-74%), specificity of 95% (95% CI 93-97%), PPV of 51% (95% CI 37-64%), NPV of 96.4% (95% CI 94.5-97.8%), and TNR of 90% (95% CI 88-93%).Attorney Docket No. 70528-701.601Table 17. Performance Characteristics of the Multi-Modal Diagnostic Workflow at the 0.54 ThresholdExample 8: Validation of a Machine Learning Algorithm for Multi-Modal Diagnostic / Biomarker Test

[0236] Using systems and methods of the present disclosure, a trained machine learning workflow was used to develop an algorithm for use in the multi-modal diagnostic assay in Example 2. In order to validate the findings, a cohort of patients with hematuria were tested using the assay, and the results compared to cystoscopy results for a confirmed positive or negative diagnosis of urothelial cancer (UC). Patients with confirmed hematuria were recruited. Eligible patients were being evaluated by flexible or rigid cystoscopy as clinically indicated, including those who were referred due to suspicious or positive imaging, and were able to provide a voided urine sample (>30 mL) and comply with study requirements. Midstream urine samples were collected from patients prior to cystoscopy or >14 days after their prior diagnostic cystoscopy procedure. Urine samples were collected on the day of (but prior to) cystoscopy inAttomey Docket No. 70528-701.601562 / 615 patients (91.4%) and at a median of 20 days before cystoscopy in 53 patients. The samples were collected and processed using the Cxbladder Urine Sampling System and analyzed by Pacific Edge Diagnostics New Zealand (PEDNZ).

[0237] An analysis cohort was recruited. A total of 702 patients with hematuria provided samples. Of the 702 patients, 87 were excluded from the analysis cohort for reasons including a prior history of bladder malignancy, prostate or renal cell carcinoma, prior genitourinary manipulation within 14 days of urine sample collection, alkylating agent-based chemotherapy, current pregnancy, or other protocol deviations. Other reasons for exclusion from the analysis cohort included: a rejection of the urine sample due to not meeting PEDNZ laboratory quality control standards, a screen failure, unconfirmed tumors (i.e., no pathological confirmation of tumor), patient consent was verified too late, or no cystoscopy was performed within 90 days. Of the 702 patients, 615 patients were included in the analysis cohort.

[0238] In order to move forward with analyzing the diagnostic test, the analysis set was compared to the cystoscopy results for a confirmed positive or negative diagnosis of UC. A confirmed positive or negative diagnosis of UC was determined according to the following criteria: patients with a pathological confirmation of UC (from biopsied or resected tissue from cystoscopy and / or ureteroscopy) were considered UC positive; patients who had no abnormalities on cystoscopy or identified from biopsied or resected tissue from cystoscopy and / or ureteroscopy were considered UC negative; patients who underwent treatment that ablated tissue without biopsy were classified as UC positive without confirmation (such patients were documented fully and excluded from analysis); and patients who were diagnosed with papillary urothelial neoplasm of low malignant potential (PUNLMP) were considered UC negative for this study, with PUNLMP considered as an alternative diagnosis for the hematuria event. Pathology-confirmed tumors were classified as high grade (HG; i.e., high grade, stage >Ta) or carcinoma in situ (Cis) tumors, or low grade (LG; i.e., low grade, Ta, unknown). AUA risk stratification was conducted retrospectively based on available data.

[0239] Next, the analysis cohort was tested by the diagnostic test to compute a score. Of the 615 patients of the analysis cohort, 28 did not yield a result from the diagnostic test. Reasons for no result include inflammation, technical replicate failure, insufficient DNA, or qPCR failure. FIG.6 shows a flow diagram of patient selection for the analysis set. In this example, the algorithm used two pre-specified score thresholds to indicate the results as either low, intermediate, or a high risk result. For scores lower than 0.15, the risk was determined to be low; for score between 0.15-0.54, the risk was determined to be intermediate; for scores greater than 0.54, the risk was high. For scores that are low-risk, the result of the diagnostic test was negative. For scores that are intermediate- or high-risk, the result of the diagnostic test was positive. Of the 587 patientsAttorney Docket No. 70528-701.601with results from the multi-modal diagnostic workflow, 171 patients (29.1%) tested Intermediate or High risk of UC, including 45 / 48 patients with confirmed tumors, and 416 (70.9%) tested Low risk of UC (i.e., ‘ruled out’). For the remaining 126 patients (i.e., false positives), central urine cytology classified 94 cases (74.6%) as benign, 19 (15.1%) as atypical, four (3.2%) as suspicious, and two (1.6%) as malignant; cytology findings were not available for seven patients.Example 9: Calculation of Performance Characteristics of a Multi-Modal Diagnostic Workflow

[0240] The multi-modal diagnostic workflow was validated using the results from Example 8. Calculations of sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), and a test-negative rate (TNR) were calculated from the 587 diagnostic test results compared with cystoscopy as described above. As shown in Table 18, the multi-modal diagnostic workflow had sensitivity of 94% (95% confidence interval 83-99%), specificity of 77% (73-80%), positive predictive value (PPV) of 26% (20-34%), negative predictive value (NPV) of 99.3% (99.3-99.9%), and a test-negative rate (TNR) of 71% (67-75%).Table 18. Performance Calculations of Multi-Modal Diagnostic WorkflowAttomey Docket No. 70528-701.601Example 10: The Multi-Modal Diagnostic Workflow Compared with AUA MH risk criteria or GH status & Hematuria Status Risk Stratification

[0241] AUA risk stratification was conducted retrospectively on the analytic set from Example 1, based on available data. Similarly, hematuria-based risk stratification was also conducted on the analytic set from Example 1. FIG. 9 shows the 2025 AUA MH risk stratification and GH status classified 598 / 615 patients (97.4%) as requiring further evaluation in this cohort, 320 with high- or intermediate-risk MH and 278 with GH. This broad classification captured all 48 UC cases but with a PPV of 8.0% (48 / 598). Notably, among patients with MH, only 17 (5.0%) were classified as low risk and 51 (15.1%) as intermediate risk by AUA criteria. Pathology-confirmed UC tumors were detected in 48 / 598 (8.0%) AUA high-risk patients, compared with 45 / 171 (26.3%) multi-modal diagnostic workflow-positive cases. This corresponds to a 3.3-fold (26.3% / 8.0%) improvement in PPV with the multi-modal diagnostic workflow over AUA risk stratification.Example 11: Analytical Validation of Six DNA SNVs of the Multi-Modal Diagnostic Workflow

[0242] A developmental dataset was gathered to conduct the analytical validation of the six DNA single-nucleotide variant (SNVs) from FGFR3 (R248C, S249F / C, G372C, Y375C) and TERT (C228T, C250T). The dataset included stabilized urine samples stored at -80 degrees C from patients with gross hematuria or microhematuria who had participated in a previous clinical validation study and patients with hematuria who participated in the STRATA: Safe Testing of Risk for Asymptomatic Microhematuria trial. The samples were blinded prior to their use in the development dataset. A total of 987 samples were included in the dataset.Linearity

[0243] The linearity of the multi-modal diagnostic workflow of Example 1 was assessed for each of the six DNA SNVs o FGFR3 (R248C, S249F / C, G372C, Y375C) and 7ERZ(C228T, C250T), with an upper limit being 2.21 L and 1.76 , respectively (where L is the mean number of target DNA molecules per droplet in droplet-digital PCR (ddPCR) using Poisson distribution). To confirm the linearity of each analytic target, a 10-point dilution curve (33,000, 6600, 1320, 264, 132, 66, 33, 16.5, 8.25, and 4.12 DNA copies per well) was performed, with at least four replicates at each concentration. Then, statistical testing was performed to determine the concentration at which the assay became non-linear for each analytic target. Data were fitted using a linear regression model, and linearity was assessed using the R2and mean squared error (MSE). The null hypothesis of this example was that the multi-modal diagnostic workflow hadAttomey Docket No. 70528-701.601valid linear regression through the whole diluted range; this null hypothesis was rejected if R2was <0.9. Then, the MSE was used to compare linearity between different analytic targets.

[0244] The multi-modal diagnostic workflow of Example 1 demonstrated linearity across all six FGFR3 and TERT SNVs with an R2value of >0.99 for all analytic targets (as seen in FIG. 5).The MSE values for individual SNVs ranged from 0.006 to 0.014. The tested range was 1 x 105to 1x101DNA copies / well, which corresponded to a maximum linear concentration of 5k. The results of this example demonstrate that the linearity of this assay mt the criteria for linearity in diagnostic testing.Analytical Sensitivity

[0245] Next, the analytical sensitivity, or the limit of detection (LOD), the multi-modal diagnostic workflow was evaluated. Analytical sensitive was defined as the lowest analyte concentration that could be consistently detected with 95% probability. The LOD for the DNA SNVs of FGFR3 and TERT was determined by the logistic regression and the three-concentrations approach.

[0246] For logistic regression, a model was fitted to explain the relationship between the dependent (positive / negative) and independent (analyte concentration) variables, and the LOD was predicted based on this fitted logistic model. For the three-concentrations approach, a fraction of positives was first calculated from highest to lowest concentration, and the concentration that matched 95% of positives (co) was the boundary of LOD.

[0247] Samples with concentrations >co were grouped into three concentration levels: ci, C2, and C3, where ci > C2 > C3 > co. If m was the number of positive samples at concentration ci, it was assumed that m ~ Poisson(kci) (where i = 1, 2, 3); the expected value E(m) = kci was the average number of positive samples at concentration ci. If pi was the probability of no samples detecting positive at concentration ci, then pi = P(m = 0) = e- ci. If c* was the LOD, by definition, 1-pi = 0.95 and pi = e- c* = 0.05; hence, c* = log 0.05 / - . If Xi is the observed number of positive samples at concentration Ci, the maximum likelihood estimator for X (A) can be computed using the multi-modal diagnostic workflow algorithm of Example 1. The three concentrations (ci, C2, and C3) were selected in multiple ways and the corresponding number of samples at each concentration changed accordingly, with both (the number of samples and concentration) having an impact on the LOD estimate. The most conservative value was used as the estimated LOD.

[0248] The predicted LOD of the multi-modal diagnostic workflow of Example 1 is shown below in Table 19. When using the logistic approach, The predicted LOD of Triage Plus for DNA detection using the logistic regression ap-proach was a mutant-to-WT DNA ratio of 1:840, 1:1200, 1:1250, and 1:970 for FGFR3 R248C, S249F / C, G372C, and Y375C, respectively, and aAttorney Docket No. 70528-701.601mutant-to-WT DNA ratio of 1:440 and 1:740 for TERT C228T and C250T, respectively (Table 1). Using the three concentra-tions approach, the predicted LOD was a mutant-to-WT DNA ratio of 1:632, 1:1220, 1:946, and 1:439 for FGFR3 R248C, S249F / C, G372C, and Y375C, respectively, and a mu-tant-to-WT DNA ratio of 1 :319 and 1 :418 for TERT C228T and C250T, respectively. There was no significant difference between the logistic regression and three concentrations ap-proach for LOD prediction; however, the logistic regression approach was preferred for all SNVs except FGFR3 G372C, for which the logistic regression approach did not work well as there were too few negative samples.Table. 19. Limit of Detections of 6 DNA SNVsAnalytical Specificity

[0249] Next, the analytical specificity of the multi-modal diagnostic workflow was evaluated. Analytical specificity was defined as the ability of the multi-modal diagnostic workflow to detect FGFR3 and TERT mutant DNA in the presence of potentially interfering substances, which may have been carried over from the extraction reagent or present in the patient’s urine sample. The effect on analytical specificity of the assay for both sample- and process-derived interfering substances was assessed. The sample-derived substances assessed were red blood cells (RBCs; 8 * 105, 4 x io6, 2 * 107, and 1 x io8cells / mL), bacteria (1xio6cells / mL Escherichia coli), yeast (1 x 104colony-forming units [cfu] / mL), urea (60 mg / mL), glucose (0.5 mg / mL), and protein (1.25, 2.50, 5.00, and 10.00 mg / mL serum albumin). The selected amount of each substance was mixed with eight high- and low-concentration FGFR3 and TERTAttorney Docket No. 70528-701.601extraction controls. The contaminated controls were then extracted and compared with high- and low-concentration controls without interfering substances. Samples within the expected level of gene variance and where the multi-modal diagnostic workflow score was within the 95% confidence interval (CI) for the control were considered acceptable. The difference was calculated by subtracting the mean score for the control sample from that of each contaminated sample. The corresponding p-value was calculated using a two-sided t test.

[0250] The results of the experiment for sample-derived substances are shown below in Table 20. For sample-derived substances, the results show that RBCs had no significant effect on extraction of HECs or LECs of FGFR3 and TERT at RBC concentrations below 4 / | 06cells / mL. At RBC concentrations of 2 * 107cells / mL, the FGFR3 and TERT mutant count per well was increased for HECs and decreased for LECs. At these RBC concentrations, all controls returned a positive multi-modal diagnostic workflow result. At RBC concentrations of 1 x 108cells / mL, there was significant impact on extraction efficiency; however, there was no loss of sensitivity as all HECs returned a positive multi-modal diagnostic workflow result. LECs had some loss of sensitivity at RBC concentrations of 1 x 108cells / mL, with the potential to return a false-negative test result.Table 20. Analytical specificity of the Multi-Modal Diagnostic Workflow for FGFR3 and TERT DNA control samples when mixed with sample-derived interfering substancesAttorney Docket No. 70528-701.601

[0251] The process-derived substances assessed were ethanol (4%), MagMAX wash buffer (2%), Cxbladder stabilizing reagent (1%), MagMAX magnetic beads (5%), and acetone (10%). The substance percentage was per 64 pL of elution volume and each substance was mixed with high- and low-concentrations of FGFR3 and TERT extraction controls as 12 replicates and compared with controls (without potentially interfering substances) to determine if the samples were affected by the process-derived contaminants.

[0252] For process-derived substances, the results show that all reagents showed a slight increase in assay performance, but were within the expected level of variation. The substances that most affected the assay in this example, MagMAX wash buffer and Cxbladder stabilizing reagent, are the first and second reagent in the multistep extraction and purification; therefore, the risk of these reagents being present at the elution step at concentrations that may affect the assay is very low. The substances that are closest to the elution step (acetone and MagMAX magnetic beads) did not significantly impact the assay at the high concentrations tested.Attorney Docket No. 70528-701.601Analytical Accuracy

[0253] Next, the analytical accuracy of the multi-modal diagnostic workflow of Example 1 was evaluated. Multiplex ddPCR was considered the most accurate method to define absolute quantitation in Triage Plus. To validate the analytical accuracy of Triage Plus, ddPCR of the SNVs (AFGFR3 (R248C, S249F / C, G372C, and Y375C) and TEAT(C228T, C250T) as single analytes were compared with combined analyte samples (of mutant plus wild type (WT) DNA). A combination of TERT C228T and C250T mutant DNA was also assessed.

[0254] Then, mutant DNA — alone and combined with WT DNA — were manufactured in parallel to contain equivalent concentrations of the target DNA. High-extraction controls (HECs) had a DNA concentration of ~1 x 106copies / pL and low-extraction controls (LECs) had a DNA concentration of ~1 x 104copies / pL. Each mutant was combined with WT DNA at a high (1:10) and low (1 :200) mutant:WT ratio for the HECs and LECs, respectively. Mutant DNA samples were compared with mutant + WT DNA samples for each SNV of FGFR3 and TERT in the multiplex ddPCR assays.

[0255] Accuracy was determined as the percentage of the expected mutant DNA concentration compared with the mutant + WT DNA sample. Quantification of mutant DNA within the combined sample was then reviewed against its corresponding 95% CI.

[0256] The results of the experiment are shown below in Table 21. The absolute quantification of FGFR3 mutant DNA versus mutant + WT DNA were within the 95% CI the expected intraplate variance for HECs of R248C and S249F / C and LECs of S249F / C and G372C (Table 3). However, higher intra-plate variance was observed for the LEC of R248C, HEC of G372C, and the HEC and LEC of Y373. TERT mutant DNA versus mutant + WT DNA quantification showed acceptable intra-plate variance for HECs of C228T and C228T + C250T, but higher acceptable intra-plate variance for all LECs and the HEC of C250T.Table 21. Results of Analytical Accuracy for the Six SNVsAttorney Docket No. 70528-701.601Analytical Precision

[0257] Next, analytical precision of the multi-modal diagnostic workflow assay was evaluated. Analytical precision was defined as reproducibility within a single run or between separate runs for replicate samples. Variance in the ddPCR assay was assessed using HECs and LECs for each FGFR3 and TERT SNV.

[0258] A fitted linear random effects model was used to estimate intra-assay and inter-assay variance for the mutant fraction, presented as standard deviation (SD) and 95% CI; the coefficient of variation (CV%) was also calculat-ed. Inter-assay, intra-assay, and total variance for the mutant fraction were assessed by four operators over >60 days by reviewing 46 ddPCR plate controls split over 22 plates (per FGFR3 SNV) or 48 ddPCR plate controls over 23 plates (per TERT SNV).

[0259] The lot-to-lot reagent variation was also assessed by testing three independent manufacture lots of MagMAX magnetic beads (BD2206339, BD2206338, and BD2302342) and MagMAX wash buffer (WB2405057, WB2405056, and WB2312054) for both FGFR3 and TERT SNVs using HECs and LECs. Each manufacture lot was run in parallel on a single plate with at least eight replicates of each control for each lot of reagents.

[0260] The results for the analytical precision of the multi-modal diagnostic workflow are shown in Table 22 below. Inter-assay variability showed mutant fraction CV% of <4.33% for HECs, with mutant fraction CV%s of 1.35-4.33% for FGFR3 and 1.01-2.11% for TERT, however, mutant fraction variance was slightly higher for LECs (mutant fraction CV%s of 4.62-Attorney Docket No. 70528-701.6016.97% and 3.42-6.86%, respectively). Based on a maximum mutant fraction CV% of 6.97% for FGFR3 C228T and 6.86% for TERT C228T + C250T, the inter-assay variability was deemed acceptable. For intra-assay variability, HECs had mutant fraction CV%s of <3.54% for FGFR3 and <2.95% for TERT. LECs had higher mutant fraction CV%s than HECs (11.56-15.90% for FGFR3 and 11.79-20.27% for TERT). Based on a maximum mutant fraction CV% of 15.90% for FGFR3 G372C and 20.27% for TERT C228T + C250T, the intra-assay variability was deemed acceptable. Total assay mutant fraction CV% for FGFR3 and TERT was <5.59% for HECs, and 12.45-17.24% and 12.27-21.40%, respectively, for LECs. The maximum total mutant fraction variance was considered excellent for high-extraction controls and acceptable for low-extraction controls.Table 22. Analytical Precision of the Multi-Modal Diagnostic WorkflowAttorney Docket No. 70528-701.601

[0261] As shown by Table 23, Lot-to-lot reagent variance was low for FGFR3 HECs (mutant fraction CV%s 1.14-2.53%), and within an acceptable range (CV%s 13.10-17.64%) for LECs. Similarly with TERT, lot-to-lot variance was low (mutant fractions CV%s 0.84-1.74%) for HECs, but was slightly higher for LECs (CV%s 11.46-32.15%).Table 23. Lot-to-Lot Reagent VarianceExtraction Efficiency

[0262] Next, the extraction efficiency was evaluated to confirm that the multi-modal diagnostic workflow of Example 1 was not biased for extraction of either FGFR3 or TERT mutant DNA,Attorney Docket No. 70528-701.601synthetic urine samples were prepared for each SNV high (DNA concentration of ~1 x 106 copies / pL and mutant:WT ratio of 1:10) and low (DNA concentration of ~1 x 104 copies / pL and mutant:WT ratio of 1:200) mutant:WT ratios. The extraction efficiency was tested by processing eight replicates of each synthetic urine sample from extraction to ddPCR. These results were then compared with the expected mutant fraction and copy number. The absolute quantity of the extracted sample was used to define the extraction efficiency of DNA.

[0263] The results of the evaluation are shown in Table 24. The extraction efficiency of the multi-modal diagnostic workflow for FGFR3 and TERT SNVs was lower in HECs versus LECs. The FGFR3 samples had a mean extraction efficiency of 72.9% for high mutant-to-WT ratio controls and 88.0% for low mutant-to-WT ratio controls. The TERT samples had a mean extraction efficiency of 83.5% for high mutant-to-WT ratio controls and 95.5% for low mutant-to-WT ratio controls. No difference was observed between extracted samples and input controls, confirming that the extraction method did not introduce sampling bias.Table 24. Extraction Efficiency of the Multi-Modal Diagnostic WorkflowAttorney Docket No. 70528-701.601Inter-Laboratory Evaluation

[0264] An inter-laboratory comparison of the multi-modal diagnostic workflow of Example 1 was conducted between the NZ laboratory (PEDNZ) and the US laboratory (PEDUSA). A random sample set from the PEDNZ validation was used to confirm the reproducibility of the multi-modal diagnostic workflow at PEDUSA. Acceptable variability was defined as achieving >80% concordance for all clinical results. The inter-laboratory comparison between PEDNZ and PEDUSA was based on data from 33 samples. The results showed there was 87.9% concordance in clinical results for the multi-modal diagnostic workflow between the two laboratories.Example 12: Calculation of Performance Characteristics of the Multi-Modal Diagnostic Workflow with a Pre-Specified Thresholds of 0.15 and 0.54

[0265] The multi-modal diagnostic workflow was tested using a different pre-specified threshold from Examples 8-1 E The analytic set from Example 11 was tested using thresholds of 0.15 and 0.54. As shown in Table 25, under the 0.54 threshold for the multi-modal diagnostic workflow, 63 / 987 patients (6.4%) had a high risk result, and 47 / 78 tumors (60.3%) were detected. Under the 0.15 threshold for the multi-modal diagnostic workflow, 157 / 987 patients (15.9%) had a high risk result, and 73 / 78 tumors (93.6%) were detected. When the lower and upper score thresholds for the multi-modal diagnostic workflow were used to define the probability of UC, the actual incidence of UC (confirmed by pathology) was 0.6% in low-probability samples (score < 0.15; 84.1% of samples), 27.7% in intermediate-probability samples (score > 0.15 to <0.54; 9.5% of samples), and 74.6% in high-probability samples (score > 0.54; 6.4% of samples).Table 25. Confusion Matrix of the Multi-Modal Diagnostic Workflow at the Upper and Lower Thresholds Versus Confirmed Tumor StatusAttomey Docket No. 70528-701.601

[0266] Next, calculations of performance characteristics were calculated for the multi-modal diagnostic workflow at the 0.54 threshold. Calculations of sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), and a test-negative rate (TNR) were calculated from the 587 diagnostic test results compared with cystoscopy as described above. Under the 0.15 threshold, the multi-modal diagnostic workflow gave a sensitivity of 93.6%, a specificity of 90.8%, a PPV of 46.5%, an NPV of 99.4%, and a TNR of 84.1%. Using the upper score threshold (0.54), the multi-modal diagnostic workflow had higher specificity (98.2%) and PPV (74.6%).

[0267] While preferred embodiments of the present inventive concepts have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the inventive concepts. It should be understood that various alternatives to the embodiments of the inventive concepts described herein may be employed in practicing the inventive concepts. It is intended that the following claims define the scope of the inventive concepts and that methods and structures within the scope of these claims and their equivalents be covered thereby.

Claims

1. Attorney Docket No. 70528-701.601CLAIMS1. A method for analyzing a sample, the method comprising:a) assaying the sample or a first portion of the sample obtained from a subject that has bladder cancer or is suspected of having bladder cancer with a genotyping assay to detect one or more genotypes at one or more polymorphisms, thereby generating a first data set, wherein the one or more polymorphisms comprise rsl21913482, rsl21913483, rsl21913479, rsl21913485, rsl242535815, rsl561215364, or a polymorphism in linkage disequilibrium therewith as determined by a coefficient of determination R2of at least 0.85, or any combination thereof; b) assaying the sample or a second portion of the sample with a transcriptomic assay to determine a quantitative measure of one or more transcriptomic markers, thereby generating a second data set, wherein the one or more transcriptomic markers comprise Midkine (MDK), Cyclin Dependent Kinase 1 (CDK1), Insulin Like Growth Factor Binding Protein 5 (IGFBP5), Homeobox Al 3 (HOXA13), or C-X-C Motif Chemokine Receptor 2 (CXCR2), or any combination thereof; andc) analyzing the first data set and the second data set using a trained machine learning algorithm to determine an output indicative of the subject as having a degree of cancer risk.

2. The method of claim 1, further comprising generating an electronic report comprising the output indicative of the subject as having the degree of cancer risk.

3. The method of claim 1, wherein the subject has or is suspected of having hematuria.

4. The method of claim 1, wherein the one or more genotypes are heterozygous for a risk allele at the one or more polymorphisms.

5. The method of claim 1, wherein the one or more polymorphisms comprise two or more of rsl21913482, rsl21913483, rsl21913479, rsl21913485, rsl242535815, rsl561215364, or a polymorphism in linkage disequilibrium therewith as determined by a coefficient of determination R2of at least 0.85, or any combination thereof.

6. The method of claim 1, wherein the one or more polymorphisms comprise rsl21913482, rsl21913483, rsl21913479, rsl21913485, rsl242535815, and rsl561215364.

7. The method of claim 1, wherein the one or more transcriptomic markers comprise MDK, CDK1, IGFBP5, HOXA13, and CXCR2.Attomey Docket No. 70528-701.6018. The method of claim 1, wherein:the one or more transcriptomic markers comprise MDK, CDK1, IGFBP5, HOXA13, and CXCR2; andthe one or more polymorphisms comprise rsl21913482, rsl21913483, rsl21913479, rsl21913485, rsl242535815, and rsl561215364.

9. The method of claim 1, wherein the quantitative measure comprises a level or an amount of the one or more transcriptomic markers.

10. The method of claim 1, wherein the trained machine learning algorithm comprises a trained machine learning classifier, wherein the degree of cancer risk comprises a categorical cancer risk selected from among a plurality of distinct categorical cancer risks, and wherein determining the output comprises assigning a classification of the subject as having the categorical cancer risk selected from among the plurality of distinct categorical cancer risks.

11. The method of claim 10, wherein the plurality of distinct categorical cancer risks comprises a high cancer risk, an intermediate cancer risk, or a low cancer risk, or any combination thereof.

12. The method of claim 11, wherein the plurality of distinct categorical cancer risks are determined with reference to thresholds based at least in part on a highest calculated Youden index of a Receiver Operating Characteristic (ROC) curve for the trained machine learning algorithm.

13. The method of claim 11, wherein the assigned classification is the high cancer risk or the intermediate cancer risk, and wherein the method further comprises administering a treatment to the subject capable of treating the bladder cancer.

14. The method of claim 13, wherein the treatment comprises a surgical resection, a chemotherapy, an immunotherapy, a radiation therapy, a targeted therapy, or a combination thereof.

15. The method of claim 10, wherein the classification has a specificity of at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, or at least about 95%.Attomey Docket No. 70528-701.60116. The method of claim 10, wherein the classification has a sensitivity of at least about 85%, at least about 90%, or at least about 95%.

17. The method of claim 10, wherein the classification has a negative predictive value (NPV) of at least about 95%.

18. The method of claim 10, wherein the classification has a negative predictive value (NPV) of at least about 96.5%.

19. The method of claim 10, wherein the classification a positive predictive value (PPV) of at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, or at least about 95%.

20. The method of claim 1, wherein the trained machine learning algorithm comprises a nonparametric model.

21. The method of claim 20, wherein the trained machine learning algorithm comprises a boosted regression tree (BRT), a Bayesian Additive Regression Tree (BART), a random forest (RF) algorithm, a boosted tree algorithm, a gradient boosting machine (GBM), a Bayesian Model Averaging (BMA) with a decision tree, a classification and regression tree (CART) model, a support vector machine (SVM), a Linear Discriminate Analysis (LDA), a Logistic Regression (LogReg), a K-nearest 5 neighbors (Kn5n), a partition tree classifier (TREE), or a partially fixed BART, or any combination thereof.

22. The method of claim 21, wherein the trained machine learning algorithm comprises the BART, the RF, the GBM, or a Bayesian CART.

23. The method of claim 1, wherein the analyzing the first data set and the second data set using the trained machine learning algorithm is performed with a single trained machine learning algorithm.

24. The method of claim 1, wherein the trained machine learning algorithm is an integrated algorithm or an ensemble machine learning algorithm comprising two or more algorithms.Attomey Docket No. 70528-701.60125. The method of claim 1, wherein the trained machine learning algorithm comprises an inverse logit function.

26. The method of claim 25, wherein the inverse logit function is an inverse logit of a second-order polynomial.

27. The method of claim 26, wherein the second-order polynomial comprises coefficients obtained by fitting a logistic regression model.

28. The method of claim 25, wherein the inverse logit function determines an estimated probability of the subject having cancer.

29. The method of claim 28, further comprising risk stratifying the subject based at least in part on thresholding the estimated probability of the subject having cancer.

30. The method of claim 1, wherein the degree of cancer risk comprises an estimated probability of cancer.

31. The method of claim 1, wherein the sample is a urine sample.

32. The method of claim 31, further comprising collecting the urine sample from the subject.

33. The method of claim 1, wherein the transcriptomic assay comprises a droplet digital polymerase chain reaction (PCR).

34. The method of claim 1, wherein the genotyping assay comprises a quantitative reverse transcription analysis.

35. The method of claim 1, wherein the bladder cancer is urothelial bladder cancer.

36. The method of claim 1, further comprising training the trained machine learning algorithm prior to (a).Attomey Docket No. 70528-701.60137. The method of claim 1, wherein the training comprises:receiving a genotype data and transcriptomic data from a plurality of training samples comprising case samples obtained from subjects with the bladder cancer, and control samples obtained from subjects without the bladder cancer;creating a first training set comprising covariate data corresponding to the genotype data and the transcriptomic data; andtraining a machine learning algorithm with the covariate data to determine an output indicative of the subject as having a degree of cancer risk.

38. The method of claim 1, wherein the trained machine learning algorithm is trained with a training set that is independent of the sample.

39. The method of claim 1, wherein the trained machine learning algorithm has a performance characterized by a receiver operating characteristic (ROC) curve having an average or median area under the curve (AUC) of at least 0.85.

40. The method of claim 1, wherein the trained machine learning algorithm has a performance characteristic comprising a specificity of at least about 70% when the trained machine learning algorithm is trained on bladder cancer samples and non-cancer samples.

41. The method of claim 1, wherein the trained machine learning algorithm has a performance characteristic comprising a sensitivity of at least about 85% when the trained machine learning algorithm is trained on bladder cancer samples and non-cancer samples.

42. The method of claim 1, wherein the trained machine learning algorithm has a performance characteristic comprising a negative predictive value (NPV) of at least about 95% when the trained machine learning algorithm is trained on bladder cancer samples and non-cancer samples.

43. The method of claim 1, wherein the trained machine learning algorithm has a performance characteristic comprising a negative predictive value (NPV) of at least about 96.5% when the trained machine learning algorithm is trained on bladder cancer samples and non-cancer samples.Attomey Docket No. 70528-701.60144. The method of claim 1, wherein the trained machine learning algorithm has a performance characteristic comprising a positive predictive value (PPV) of at least about 15% when the trained machine learning algorithm is trained on bladder cancer samples and noncancer samples.

45. The method of claim 11, further comprising identifying the subject as not having the bladder cancer or having a low risk of developing the bladder cancer, based at least in part on the assigned classification being the low cancer risk, thereby avoiding having to perform a cystoscopy on the subject.

46. The method of claim 11, further comprising performing a cystoscopy on the subject to confirm a diagnosis of the bladder cancer, based at least in part on the assigned classification being the high cancer risk.

47. The method of claim 11, further comprising calculating a risk score corresponding to the high cancer risk, wherein the risk score exceeds the threshold of 0.54.

48. The method of claim 11, further comprising calculating a risk score corresponding to the intermediate cancer risk, wherein the risk score is from 0.15 to 0.54.

49. The method of claim 11, further comprising calculating a risk score corresponding to the low cancer risk, wherein the risk score is below 0.15.

50. A computer-implemented system comprising a computing device comprising at least one processor, an operating system configured to perform executable instructions, a memories storing machine-executable code that, when executed, causes the at least one processor to: a) receive a first data set obtained from assaying a sample or a first portion of the sample obtained from a subject that has bladder cancer or is suspected of having bladder cancer with a genotyping assay to detect one or more genotypes at one or more polymorphisms, wherein the one or more polymorphisms comprise rsl21913482, rsl21913483, rsl21913479, rsl21913485, rsl242535815, rsl561215364, or a polymorphism in linkage disequilibrium therewith as determined by a coefficient of determination R2of at least 0.85, or any combination thereof; b) receive a second data set obtained from assaying the sample or a second portion of the sample obtained from the subject that has bladder cancer or is suspected of having bladder cancer with a transcriptomic assay to determine a quantitative measure of one or more transcriptomic markers to generate the second data set, wherein the one or more transcriptomicAttomey Docket No. 70528-701.601markers comprise Midkine (MDK), Cyclin Dependent Kinase 1 (CDK1), Insulin Like Growth Factor Binding Protein 5 (IGFBP5), Homeobox A13 (HOXA13), or C-X-C Motif Chemokine Receptor 2 (CXCR2), or any combination thereof; andc) analyze the first data set and the second data set using a trained machine learning algorithm to determine an output indicative of the subject as having a degree of cancer risk.

51. The system of claim 50, wherein the one or more genotypes are heterozygous for a risk allele at the one or more polymorphisms.

52. The system of claim 51, wherein the one or more polymorphisms comprise two or more of rsl21913482, rsl21913483, rsl21913479, rsl21913485, rsl242535815, rsl561215364, or a polymorphism in linkage disequilibrium therewith as determined by a coefficient of determination R2of at least 0.85, or any combination thereof.

53. The system of claim 50, wherein the one or more polymorphisms comprise rsl21913482, rsl21913483, rsl21913479, rsl21913485, rsl242535815, and rsl561215364.

54. The system of claim 50, wherein the one or more transcriptomic markers comprise MDK, CDK1, IGFBP5, HOXA13, and CXCR2.

55. The system of claim 50, wherein:the one or more transcriptomic markers comprises MDK, CDK1, IGFBP5, HOXA13, and CXCR2; andthe one or more polymorphisms comprises rsl21913482, rsl21913483, rsl21913479, rsl21913485, rsl242535815, and rsl561215364.

56. The system of claim 50, wherein the quantitative measure comprises a level or an amount of the one or more transcriptomic markers.

57. The system of claim 50, wherein the trained machine learning algorithm comprises a trained machine learning classifier, wherein the degree of cancer risk comprises a categorical cancer risk selected from among a plurality of distinct categorical cancer risks, and wherein determining the output comprises assigning a classification of the subject as having the categorical cancer risk selected from among the plurality of distinct categorical cancer risks.Attomey Docket No. 70528-701.60158. The system of claim 57, wherein the plurality of distinct categorical cancer risks comprises a high cancer risk, an intermediate cancer risk, or a low cancer risk, or any combination thereof.

59. The system of claim 58, wherein the plurality of distinct categorical cancer risks are determined with reference to thresholds based at least in part on a highest calculated Youden index of a Receiver Operating Characteristic (ROC) curve for the trained machine learning algorithm.

60. The system of claim 58, wherein the assigned classification is the high cancer risk or the intermediate cancer risk, and wherein the machine-executable code, when executed, causes the at least one processor to output a report comprising a recommended treatment of the bladder cancer.

61. The system of claim 60, wherein the treatment comprises a surgical resection, a chemotherapy, an immunotherapy, a radiation therapy, a targeted therapy, or a combination thereof.

62. The system of claim 50, wherein the trained machine learning algorithm has a performance characterized by a receiver operating characteristic (ROC) curve having an average or median area under the curve (AUC) of at least about 0.85.

63. The system of claim 50, wherein the trained machine learning algorithm has a performance characterized by a specificity of at least about 70%.

64. The system of claim 58, wherein the sample is classified as a high cancer risk, an intermediate cancer risk, or a low cancer risk at a sensitivity of at least about 85%.

65. The system of claim 58, wherein the sample is classified as a high cancer risk, an intermediate cancer risk, or a low cancer risk at a negative predictive value (NPV) of at least about 96.5%.

66. The system of claim 58, wherein the sample is classified as a high cancer risk, an intermediate cancer risk, or a low cancer risk at a positive predictive value (PPV) of at least about 15%.Attomey Docket No. 70528-701.60167. The system of claim 50, wherein the subject has or is suspected of having hematuria.

68. The system of claim 50, wherein the trained machine learning algorithm comprises a non-parametric model.

69. The system of claim 68, wherein the trained machine learning algorithm comprises a boosted regression tree (BRT), a Bayesian Additive Regression Tree (BART), a random forest (RF) algorithm, a boosted tree algorithm, a gradient boosting machine (GBM), a Bayesian Model Averaging (BMA) with a decision tree, a classification and regression tree (CART) model, a support vector machine (SVM), a Linear Discriminate Analysis (LDA), a Logistic Regression (LogReg), a K-nearest 5 neighbors (Kn5n), a partition tree classifier (TREE), or a partially fixed BART, or any combination thereof.

70. The system of claim 69, wherein the trained machine learning algorithm comprises the BART, the RF, the GBM, or a Bayesian CART.

71. The system of claim 50, wherein the analyzing the first data set and the second data set using the trained machine learning algorithm is performed with a single trained machine learning algorithm.

72. The system of claim 50, wherein the trained machine learning algorithm is an integrated algorithm comprising two or more algorithms.

73. The system of claim 50, wherein the trained machine learning algorithm comprises an inverse logit function.

74. The system of claim 73, wherein the inverse logit function is an inverse logit of a second-order polynomial.

75. The system of claim 74, wherein the second-order polynomial comprises coefficients obtained by fitting a logistic regression model.

76. The system of claim 73, wherein the inverse logit function determines an estimated probability of the subject having cancer.Attomey Docket No. 70528-701.60177. The system of claim 76, wherein the machine-executable code, when executed, causes the at least one processor to risk stratifying the subject based at least in part on thresholding the estimated probability of the subject having cancer.

78. The system of claim 50, wherein the degree of cancer risk comprises an estimated probability of cancer.

79. The system of claim 50, wherein the sample is a urine sample.

80. The system of claim 50, wherein the transcriptomic assay comprises a droplet digital polymerase chain reaction (PCR).

81. The system of claim 50, wherein the genotyping assay comprises a quantitative reverse transcription analysis.

82. The system of claim 50, wherein the bladder cancer is urothelial bladder cancer.

83. The system of claim 50, wherein the trained machine learning algorithm comprises: a genotype data and transcriptomic data from a plurality of training samples comprising case samples obtained from subjects with the bladder cancer, and control samples obtained from subjects without the bladder cancer; anda first training set comprising covariate data corresponding to the genotype data and the transcriptomic data,wherein the machine learning algorithm is configured to be trained by covariate data to determine an output indicative of the subject as having a degree of cancer risk.

84. The system of claim 50, wherein the trained machine learning algorithm is trained with a training set that is independent of the sample.

85. The system of claim 50, wherein the trained machine learning algorithm has a performance characterized by a receiver operating characteristic (ROC) curve having an average or median area under the curve (AUC) of at least about 0.85.

86. The system of claim 50, wherein the trained machine learning algorithm has a performance characteristic comprising a specificity of at least about 70% when the trained machine learning algorithm is trained on bladder cancer samples and non-cancer samples.Attomey Docket No. 70528-701.60187. The system of claim 50, wherein the trained machine learning algorithm has a performance characteristic comprising a sensitivity of at least about 85% when the trained machine learning algorithm is trained on bladder cancer samples and non-cancer samples.

88. The system of claim 50, wherein the trained machine learning algorithm has a performance characteristic comprising a negative predictive value (NPV) of at least about 95% when the trained machine learning algorithm is trained on bladder cancer samples and noncancer samples.

89. The system of claim 50, wherein the trained machine learning algorithm has a performance characteristic comprising a negative predictive value (NPV) of at least about 96.5% when the trained machine learning algorithm is trained on bladder cancer samples and noncancer samples.

90. The system of claim 50, wherein the trained machine learning algorithm has a performance characteristic comprising a positive predictive value (PPV) of at least 15% when the trained machine learning algorithm is trained on bladder cancer samples and non-cancer samples.

91. The system of claim 58, wherein the classification is configured to identify the subject as not having the bladder cancer or having a low risk of developing the bladder cancer based, at least in part, on the assigned classification being the low cancer risk, thereby avoiding having to perform a cystoscopy on the subject.

92. The system of claim 58, further comprising a cystoscope that is configured to be used in a cystoscopy on the subject to confirm a diagnosis of the bladder cancer, based at least in part on the assigned classification being the high or intermediate cancer risk.

93. The system of claim 58, wherein the machine-executable code, when executed, causes the at least one processor to output a report comprising a risk score corresponding to the high cancer risk, wherein the risk score exceeds the threshold of 0.54.

94. The system of claim 58, wherein the machine-executable code, when executed, causes the at least one processor to output a report comprising a risk score corresponding to the intermediate cancer risk, wherein the risk score is from 0.15 to 0.54.Attomey Docket No. 70528-701.60195. The system of claim 58, wherein the machine-executable code, when executed, causes the at least one processor to output a report comprising a risk score corresponding to the low cancer risk, wherein the risk score is below 0.15.