Methods for high throughput diagnostics and therapeutic development of agents for cancer
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2026-03-19
AI Technical Summary
Conventional cancer diagnosis and prognosis methods, such as tissue biopsies and liquid biopsies, are invasive, costly, time-consuming, and lack accuracy for early-stage detection, making them unsuitable for large-scale screening and prone to complications.
The use of extracellular matrix bodies (ECMBs) separated from a biological fluid for automated, high-throughput analysis through histochemical staining, immunohistochemical staining, and machine learning techniques to predict cancer parameters, including cancer type and stage, using microfluidic systems and prediction machine learning models.
Enables non-invasive, cost-effective, and accurate early-stage cancer detection with high sensitivity and specificity, facilitating faster, more frequent, and less risky cancer screening and monitoring compared to conventional methods.
Smart Images

Figure US2025041123_19032026_PF_FP_ABST
Abstract
Description
Attorney Docket No.: AUMI-004 / 02WO 348385-2089 METHODS FOR HIGH THROUGHPUT DIAGNOSTICS AND THERAPEUTIC DEVELOPMENT OF AGENTS FOR CANCER CROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Patent Application Serial No. 63 / 680,567 filed August 7, 2024 and U.S. Provisional Patent Application Serial No. 63 / 760,065 filed February 18, 2025, the content of each of which is incorporated herein by reference in its entirety for all purposes. TECHNICAL FIELD
[0002] The devices, systems, and methods herein relate to predicting a cancer parameter from a biological sample. BACKGROUND
[0003] Conventional methods for diagnosis and prognosis of cancer include histological staining of a tissue biopsy where, for example, a small sample of tissue is removed from a patient and stained for analysis using a microscope. For example, pathology analysis of tissue morphology may compare a biopsy sample to a healthy control to diagnose disease based on well-defined tissue characteristics or may test the biopsied tissue for biomarkers indicative of cancer. However, a tissue biopsy is a costly, time-consuming, and usually involves an invasive medical procedure which is unsuitable for screening otherwise healthy individuals, let alone a large-scale screening program. Furthermore, liquid biopsy assays lack the accuracy to detect genetic changes from tumors in earlier disease states. Accordingly, biopsies are generally performed in response to patient symptoms or other indication of disease due to risk of complications (e.g., pain, bleeding, infection, death). However, many early-stage diseases may be subclinical (e.g., without symptoms) such that the patient and / or health care provider may be unaware of an underlying pathology. Furthermore, the number of tissue biopsies that may be performed for histology may be limited since a biopsy removes the tissue being sampled. Therefore, tissue biopsies are unsuitable for early-stage screening and simply may not be practical for other implications. Accordingly, additional systems, devices, and methods for predicting disease are desirable.Attorney Docket No.: AUMI-004 / 02WO 348385-2089 SUMMARY
[0004] Described here are systems, devices, and methods useful for diagnosing, prognosing, and / or treating a subject based on extracellular matrix bodies (ECMBs) separated from a biological fluid, or indirectly from tissue or a gel, in an automated and high throughout manner. In this way, the ECMBs may be used in predicting one or more of cancer and pre-cancer, for example using histochemical staining, immunohistochemical staining and / or machine learning techniques and analysis. In general, the methods described herein for predicting cancer may comprise receiving image data corresponding to ECMBs separated from a biological fluid of a subject, generating ECMB data based on the image data, and predicting a cancer parameter based on at least the ECMB data.
[0005] In some variations, the predicted cancer parameter may be confirmed. In some variations, the ECMB data may be analyzed using a prediction machine learning model. In some variations, the cancer parameter may be confirmed based on the ECMB data analyzed using another prediction machine learning model different from the first used prediction machine learning model. In some variations, predicting the cancer parameter may include a first prediction by one of a human and a prediction machine learning model and a second prediction using an other of the human and the prediction machine learning model.
[0006] In some variations, predicting a cancer parameter may comprise a plurality of prediction methods. In some variations, predicting a cancer parameter may include a first prediction by a human (e.g., a pathologist), a second prediction using radiomics-based image parameters, and a third prediction using a prediction machine learning model.
[0007] In some variations, the ECMB data may comprise one or more morphological features. In some variations, the one or more morphological features may correspond one or more of shape, size, spatial location, intensity, abundance, color, texture characteristics, edge properties, stiffness, and other biomechanical attributes of the ECMB data.
[0008] In some variations, the method may include identifying a region of interest in the received image data including the one or more morphological features. In some variations, the region of interest is manually identified. In some variations, generating the ECMB data may include generating a plurality of image parameters from the region of interest. The imageAttorney Docket No.: AUMI-004 / 02WO 348385-2089 parameters may correspond to one or more of image intensity and image texture. In some variations, predicting the cancer parameter may include analyzing the plurality of image parameters using a prediction machine learning model. In some variations, the method may include selecting one more image parameters of the plurality of image parameters based on one or more of the morphological features using the prediction machine learning model. In some variations, confirming the predicted cancer parameter may be based on the plurality of image parameters. In some variation, a prediction machine learning model may be trained to analyze the ECMB data on unlabeled image parameter data.
[0009] In some variations, processing the ECMBs may include using a first stain and a second stain different from the first stain. Predicting the cancer parameter may be based on the ECMBs corresponding to the first stain and confirming the cancer parameter may be based on the ECMBs corresponding to the second stain. In some variations, the first stain and the second stain may comprise one or more of a protease inhibitor, a histochemical stain, a hematoxylin stain, an eosin stain, a histology stain, an alcian blue stain, a picrosirius red stain, an immunohistochemical (IHC) stain, immunofluorescent (IF) stain, a multiplex IHC or IF stain, multi-spectral imaging, a protein stain, glycomic stain, a nucleic acid stain, and chemical fixation.
[0010] In some variations, generating the ECMB data may include segmenting the received image data. In some variations, the ECMB data may be segmented using the prediction machine learning model. In some variations, a prediction machine learning model may be trained to segment the received image data based on a manually identified region of interest.
[0011] In some variations, the method may include training a prediction machine learning model on one or more of image data corresponding to ECMBs, ECMB data, and clinical data. For example, a prediction machine learning model may be trained using an electronic medical record. In some variations, the one or more morphological features may be segmented.
[0012] In some variations, a confidence value of the predicted cancer parameter may be generated. In some variations, predicting the cancer parameter may include a first prediction using a sensitivity-based prediction machine learning model and a second prediction using a specificity-based machine learning model. In some variations, the cancer parameter may be predicted as non-cancerous when the confidence value is below a predetermined threshold.Attorney Docket No.: AUMI-004 / 02WO 348385-2089
[0013] In some variations, analyzing the ECMB data using a prediction machine learning model may include a first prediction machine learning model associated with a first stain or a first biomarker and a second prediction machine learning model associated with a second stain or a second biomarker.
[0014] In some variations, the prediction machine learning model may comprise one or more of ResNet, Inception, boosting algorithms, bootstrap aggregation, random forests, decision trees, Vision Transformers, convolutional neural networks (CNNs), and combinations thereof.
[0015] In some variations, predicting the cancer parameter may include analyzing the ECMB data using at least one scoring rubric defining at least one feature set comprising one or more morphological features. In some variations, the at least one scoring rubric may define a first feature set associated with cancer and may assign the first feature set associated with cancer a first point value, and the at least one scoring rubric may define a second feature set associated with a non-cancerous sample and may assign the second feature set associated with the non- cancerous sample a second point value different from the first point value. In some variations, an erroneous sample defined by a third feature set of the at least one scoring rubric may be excluded. In some variations, the cancer parameter may be predicted based on a total point value.
[0016] In some variations, the biological fluid may be processed using one or more of microfluidic separation, chemical fixation, physical fixation, affinity chromatography, sedimentation, centrifugation, differential centrifugation, density gradient centrifugation, mesh filtration, diafiltration, tangential flow filtration, membrane filtration, elutriation, affinity-based capture, precipitation, immuno-affinity capture, tag-based affinity capture, protein A / G / L affinity capture, enzyme or receptor-ligand affinity capture, lectin-based affinity capture, carbohydrate and sugar-based affinity capture, capture by carbohydrate-binding modules, synthetic glycopolymer affinity capture, hybridization-based capture, aptamer-based affinity capture, affinity capture based on peptide nucleic acid probes, affinity capture based on protein- nucleic acid interactions, magnetic bead capture, ultrasonic capture, size exclusion chromatography, ion exchange chromatography, hydrophobic interaction chromatography, electrophoresis, dialysis, flow cytometry, field-flow fractionation, AC electrokinetics, and embedding. In some variations, processing the biological fluid may use a microfluidic chipAttorney Docket No.: AUMI-004 / 02WO 348385-2089 comprising an inlet reservoir, at least one channel, at least one obstruction configured to restrict fluid flow, and an outlet reservoir.
[0017] In some variations, processing the separated ECMBS may include applying to the separated ECMBs one or more of a protease inhibitor, a histochemical stain, a hematoxylin stain, an eosin stain, a histology stain, an alcian blue stain, a picrosirius red stain, an immunohistochemical (IHC) stain, a multiplex IHC stain, multi-spectral imaging, a protein stain, a nucleic acid stain, and a chemical fixation. In some variations, the ECMB data may comprise the expression levels of one or more biomarkers selected from the group consisting of Fibronectin, Tetranectin, Thrombospondin, Galectin-3 Binding Protein (3BP), Talin, Rab27B, Zyxin, DAPI, CD5L, Afamin, Carbonic Anhydrase (CA1), INF2, Clusterin, Victronectin, Gelsolin, S100A9, CD5L, and combinations thereof.
[0018] In some variations, the cancer parameter may comprise one or more of pre-cancer, a cancer type, and a cancer stage. In some variations, the pre-cancer may comprise one or more of pre-cancerous colon or rectal polyps, pre-cancer lesions, and advanced adenoma. In some variations, the cancer type may comprise one or more of non-cancer, colorectal cancer, lung cancer, pancreatic cancer, prostate cancer, esophageal cancer, liver cancer, ovarian cancer, kidney cancer, melanoma, gastric cancer, and breast cancer.
[0019] In some variations, the ECMB data may comprise at least one morphological feature of the ECMBs selected from the group consisting of a flake, a punctum, a decorator, and combinations thereof. In some variations, the flake may comprise an eosinophilic structure having an area of between about 5 μm2and about 450 μm2, the punctum may comprise a circular structure having an area less than about 15 μm2, and the decorator may comprise a circular structure having an area less than about 7 μm2. In some variations, at least one morphological feature of the ECMBs may be segmented. In some variations, the segmenting may comprise inputting the ECMB data into the prediction machine learning model. In some variations, the separated ECMBs may comprise an isolate fraction. In some variations, the biological fluid may comprise one or more of whole blood, blood plasma, and blood serum.
[0020] In some variations, an electronic medical record may be updated based on the cancer parameter prediction. In some variations a report for one or more stakeholders may be generated,Attorney Docket No.: AUMI-004 / 02WO 348385-2089 the report comprising the cancer parameter prediction. In some variations, the subject may be a human subject. In some variations, the subject may be a non-human subject. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] FIG.1 depicts a flowchart representation of an illustrative variation of predicting cancer based on ECMBs separated from a biological fluid.
[0022] FIG.2A are images of an illustrative variation of ECMBs corresponding to non- cancerous control plasma on a microfluidic chip.
[0023] FIG.2B are images of an illustrative variation of ECMBs corresponding to pancreatic cancer plasma on a microfluidic chip.
[0024] FIG.2C are images of an illustrative variation of ECMBs corresponding to colorectal cancer plasma on a microfluidic chip.
[0025] FIG.2D are images of an illustrative variation of ECMBs corresponding to prostate cancer plasma on a microfluidic chip.
[0026] FIG.2E are images of an illustrative variation of ECMBs corresponding to lung cancer plasma on a microfluidic chip.
[0027] FIG.3A depicts a flowchart representation of an illustrative variation of a method of generating a prediction machine learning model. FIG.3B depicts a flowchart representation of an illustrative variation of a random forest machine learning model.
[0028] FIG.4 depicts a flowchart representation of an illustrative variation of generating a prediction machine learning model.
[0029] FIG.5 is a plot of an illustrative variation of a received image data set for a prediction machine learning model.
[0030] FIG.6 is a flowchart representation of an illustrative variation of ensemble learning techniques for training a prediction machine learning model.Attorney Docket No.: AUMI-004 / 02WO 348385-2089
[0031] FIG.7 are equations corresponding to an illustrative variation for prediction machine learning model evaluation.
[0032] FIGS.8A-8D are images of an illustrative variations of flakes.
[0033] FIGS.9A and 9B are images of an illustrative variations of puncta.
[0034] FIGS.10A-10C are images of an illustrative variations of decorators.
[0035] FIG.11A is a table of a scoring rubric of an illustrative variation of a method of predicting cancer. FIGS.11B and 11C are images of an illustrative variations of ECMB morphologies.
[0036] FIG.12 is a block diagram of an illustrative variation of a system.
[0037] FIGS.13A and 13B depict perspective views of an illustrative variation of a system.
[0038] FIG.14 depicts a plan view of an illustrative variation of a microfluidic chip.
[0039] FIG.15A is a density plot of an illustrative variation of a gel plot for a control isolate fraction and a stage II colorectal cancer isolate fraction. FIG.15B is a density plot of an illustrative variation of a gel plot for a control isolate fraction and a stage III colorectal cancer isolate fraction.
[0040] FIG.15C are images of an illustrative gel plot for a control isolate fraction, a stage II colorectal cancer isolate fraction, and a stage III colorectal cancer isolate fraction.
[0041] FIG.16A is a binary classification plot corresponding to an illustrative variation of a prediction machine learning model. FIG.16B is a confidence score table corresponding to an illustrative variation of a prediction machine learning model.
[0042] FIG.17A is a binary classification plot corresponding to an illustrative variation of a prediction machine learning model. FIG.17B is a confusion matrix for a classification model corresponding to an illustrative variation of a prediction machine learning model.
[0043] FIG.18A is a binary classification plot corresponding to an illustrative variation of a prediction machine learning model corresponding to colorectal cancer prediction. FIG.18B is aAttorney Docket No.: AUMI-004 / 02WO 348385-2089 table of sensitivity and positive predictive values (PPV) of antibodies for predicting cancer types using an illustrative variation of a prediction machine learning model. FIG.18C and 18D are confusion matrices for a classification model using antibodies corresponding to an illustrative variation of a prediction machine learning model.
[0044] FIG.19A is a binary classification plot corresponding to an illustrative variation of a prediction machine learning model. FIG.19B is a binary classification plot corresponding to an illustrative variation of a prediction machine learning model. FIG.19C is a binary classification plot corresponding to an illustrative variation of a prediction machine learning model.
[0045] FIG.20 is a binary classification plot corresponding to an illustrative variation of a prediction machine learning model corresponding to prostate cancer prediction.
[0046] FIG.21 is an error metrics table corresponding to an illustrative variation of a prediction machine learning model corresponding to colorectal cancer prediction.
[0047] FIG.22A is a confusion matrix for a classification model corresponding to an illustrative variation of a random forest machine learning model. FIG.22B is a binary classification table corresponding to an illustrative variation of a random forest machine learning model.
[0048] FIG.23A is a feature plot corresponding to a flake of an illustrative variation of a prediction machine learning model. FIG.23B is a feature plot corresponding to puncta of an illustrative variation of a prediction machine learning model. FIG.23C is a feature plot corresponding to a decorator of an illustrative variation of a prediction machine learning model. FIG.23D is a feature plot corresponding to a flake, a punctum, and a decorator of an illustrative variation of a prediction machine learning model.
[0049] FIG.24 is a binary classification table corresponding to an illustrative variation of a scoring rubric method of predicting colorectal cancer, pancreatic cancer, breast cancer, and prostate cancer.
[0050] FIG.25 is a binary classification table corresponding to an illustrative variation of a prediction machine learning model corresponding to breast cancer prediction.Attorney Docket No.: AUMI-004 / 02WO 348385-2089
[0051] FIG.26 is a binary classification table corresponding to an illustrative variation of a prediction machine learning model corresponding to pancreatic cancer prediction.
[0052] FIG.27A is a table of a scoring rubric of an illustrative variation of a method of predicting prostate cancer.
[0053] FIGS.27B-27K are images of illustrative variations of ECMB morphologies.
[0054] FIG.27L is a binary classification table corresponding to a scoring rubric method of predicting pancreatic cancer.
[0055] FIGS.28A-28G are images of illustrative variations of ECMB morphologies.
[0056] FIG.28H is a binary classification table corresponding to a scoring rubric method of predicting breast cancer.
[0057] FIG.29 is a binary classification table corresponding to an illustrative variation of a prediction machine learning model corresponding to colorectal cancer prediction.
[0058] FIG.30 is a plot corresponding to an illustrative variation of colorectal cancer prediction sensitivity of a FDA metric, trained scientist assessment, and a prediction machine learning model corresponding to colorectal cancer prediction.
[0059] FIG.31 is a binary classification table corresponding to an illustrative variation of a prediction machine learning model corresponding to prostate cancer prediction.
[0060] FIG.32 is another binary classification table corresponding to an illustrative variation of a prediction machine learning model corresponding to prostate cancer prediction.
[0061] FIG.33A is a binary classification plot corresponding to an illustrative variation of a prediction machine learning model corresponding to lung cancer prediction. FIG.33B is another binary classification plot corresponding to an illustrative variation of a prediction machine learning model corresponding to lung cancer prediction. FIG.33C is a sample classification table corresponding to an illustrative variation of a prediction machine learning model corresponding to lung cancer prediction.Attorney Docket No.: AUMI-004 / 02WO 348385-2089
[0062] FIG.34 is a binary classification plot corresponding to an illustrative variation of a prediction machine learning model corresponding to pancreatic cancer prediction.
[0063] FIG.35A is a binary classification table corresponding to an illustrative variation of a prediction machine learning model corresponding to breast cancer prediction. FIG.35B is a binary classification plot corresponding to an illustrative variation of a prediction machine learning model corresponding to breast cancer prediction.
[0064] FIG.36 is a summary binary classification table of corresponding to illustrative variations of prediction machine learning models corresponding to cancer type (colorectal cancer, prostate cancer, lung cancer, pancreatic cancer, and breast cancer) predictions.
[0065] FIG.37 are images of an illustrative variation of a cancerous and a non-cancerous ECMB sample differentiated by a morphological feature.
[0066] FIG.38 are images of an illustrative variation of a cancerous and a non-cancerous ECMB sample differentiated by a morphological feature.
[0067] FIG.39 are images of an illustrative variation of a cancerous and a non-cancerous ECMB sample differentiated by a morphological feature.
[0068] FIG.40 are images of an illustrative variation of a cancerous and a non-cancerous ECMB sample differentiated by a morphological feature.
[0069] FIG.41 are images of an illustrative variation of a cancerous and a non-cancerous ECMB sample differentiated by a morphological feature.
[0070] FIG.42A are images of an illustrative variation of an ECMB sample corresponding to a stage I cancerous fluid sample. FIG.42B are images of an illustrative variation of an ECMB sample corresponding to a non-cancerous fluid sample.
[0071] FIG.43A are images of an illustrative variation of an ECMB sample corresponding to a stage II cancerous fluid sample. FIG.43B are images of an illustrative variation of an ECMB sample corresponding to a non-cancerous fluid sample.Attorney Docket No.: AUMI-004 / 02WO 348385-2089
[0072] FIG.44A are images of an illustrative variation of an ECMB sample corresponding to a stage III cancerous fluid sample. FIG.44B are images of an illustrative variation of an ECMB sample corresponding to a non-cancerous fluid sample.
[0073] FIG.45A are images of an illustrative variation of an ECMB sample corresponding to a stage IV cancerous fluid sample. FIG.45B are images of an illustrative variation of an ECMB sample corresponding to a non-cancerous fluid sample.
[0074] FIG.46A is an intensity distribution plot corresponding to an illustrative variation of cancer samples. FIG.46B is an intensity distribution plot corresponding to an illustrative variation of non-cancerous samples.
[0075] FIG.47 is a confusion matrix for a classification model based on a morphological feature corresponding to an illustrative variation of a prediction machine learning model.
[0076] FIGS.48A-48C are scatter plots of respective image parameters based on a morphological feature corresponding to illustrative variations of cancer prediction. FIG.48D is a scatter plot of a plurality of image parameters based on a morphological feature corresponding to an illustrative variation of cancer prediction.
[0077] FIGS.49A-49C are scatter plots of respective image parameters based on a morphological feature corresponding to illustrative variations of cancer prediction for prediction of cancer based on a morphological feature. FIG.49D is a confusion matrix for a classification model based on a morphological feature corresponding to an illustrative variation of a prediction machine learning model.
[0078] FIG.50A is a table of a scoring rubric corresponding to an illustrative variation of a method of predicting lung cancer. FIGS.50B-50F are images of illustrative variations of ECMB feature sets.
[0079] FIG.51A is a table of a scoring rubric corresponding to an illustrative variation of a method of predicting non-cancerous controls. FIG.51B is images of illustrative variations of ECMB feature sets.
[0080] FIG.52 is a binary classification table corresponding to an illustrative variation of a method of predicting lung cancer.Attorney Docket No.: AUMI-004 / 02WO 348385-2089
[0081] FIG.53 is a table of a scoring rubric corresponding to an illustrative variation of a method of predicting prostate cancer.
[0082] FIG.54 is images of illustrative variations of ECMB feature sets.
[0083] FIG.55 is a binary classification table corresponding to an illustrative variation of a method of predicting prostate cancer.
[0084] FIG.56 is a table of study participant demographics corresponding to an illustrative variation of a method of predicting colorectal cancer.
[0085] FIG.57 is a table of confidence score thresholds and associated performance metrics corresponding to an illustrative variation of a prediction machine learning model.
[0086] FIG.58 is a table of confidence score thresholds and associated performance metrics corresponding to an illustrative variation of another prediction machine learning model.
[0087] FIG.59 is a plot of confidence scores corresponding to cancer parameter predictions of an illustrative variation of a prediction machine learning model associated with eosin stained samples and an illustrative variation of a prediction machine learning model associated with IF stained samples.
[0088] FIG.60 is a table of illustrative combinations of confidence score thresholds and associated performance metrics corresponding to an illustrative variation of a combination of prediction machine learning models.
[0089] FIG.61 is a table of illustrative combinations of confidence score thresholds and associated performance metrics corresponding to another illustrative variation of a combination of prediction machine learning models.
[0090] FIG.62 is table of illustrative variations of prediction machine learning models and combinations of prediction machine learning models and associated confidence score thresholds and associated performance metrics.Attorney Docket No.: AUMI-004 / 02WO 348385-2089
[0091] FIG.63 is a plot of an accuracy statistic for various combinations of confidence thresholds corresponding to an illustrative variation of a combination of prediction machine leaning models.
[0092] FIG.64 is a plot of an accuracy statistic for various combinations of confidence thresholds corresponding to another illustrative variation of a combination of prediction machine leaning models.
[0093] FIG.65 is a table of illustrative combinations of confidence score thresholds and associated performance metrics corresponding to an illustrative variation of a combination of prediction machine learning models.
[0094] FIG.66 is a table of illustrative combinations of confidence score thresholds and associated performance metrics corresponding to another illustrative variation of a combination of prediction machine learning models.
[0095] FIG.67 is a sensitivity plot of illustrative combinations of confidence thresholds corresponding to an illustrative variation of a combination of prediction machine leaning models.
[0096] FIG.68 is a sensitivity plot of illustrative combinations of confidence thresholds corresponding to another illustrative variation of a combination of prediction machine leaning models.
[0097] FIG.69 is a specificity plot of illustrative combinations of confidence thresholds corresponding to an illustrative variation of a combination of prediction machine leaning models.
[0098] FIG.70 is a specificity plot of illustrative combinations of confidence thresholds corresponding to another illustrative variation of a combination of prediction machine leaning models.
[0099] FIG.71 are images of illustrative variations of regions of interest of a cancerous fluid sample.Attorney Docket No.: AUMI-004 / 02WO 348385-2089
[0100] FIG.72 are images of other illustrative variations of regions of interest of a cancerous fluid sample.
[0101] FIG.73 are images of other illustrative variations of regions of interest of a cancerous fluid sample.
[0102] FIG.74 are images of illustrative variations of regions of interest of a non-cancerous fluid sample.
[0103] FIG.75 are images of other illustrative variations of regions of interest of a non- cancerous fluid sample.
[0104] FIG.76 are images of an illustrative variation of a cancerous ECMB sample comparing human identified regions of interest and prediction machine learning model identified regions of interest.
[0105] FIG.77 are images of another illustrative variation of a cancerous ECMB sample comparing human identified regions of interest and prediction machine learning model identified regions of interest.
[0106] FIG.78 are images of another illustrative variation of a cancerous ECMB sample comparing human identified regions of interest and prediction machine learning model identified regions of interest.
[0107] FIG.79 are images of an illustrative variation of a non-cancerous ECMB sample comparing human identified regions of interest and prediction machine learning model identified regions of interest.
[0108] FIG.80 are images of illustrative variations of a non-cancerous ECMB sample comparing human identified regions of interest and prediction machine learning model identified regions of interest.
[0109] FIG.81 is a correlation metrics table of illustrative variations of a human prediction method and prediction machine learning model.Attorney Docket No.: AUMI-004 / 02WO 348385-2089
[0110] FIG.82 are images of illustrative variations of a cancerous ECMB sample comparing human identified regions of interest and prediction machine learning model identified regions of interest.
[0111] FIG.83 are images of illustrative variations of a non-cancerous ECMB sample comparing human identified regions of interest and prediction machine learning model identified regions of interest.
[0112] FIG.84 is a plot of an agreement statistic comparing illustrative variations of a human prediction method and prediction machine learning model.
[0113] FIG.85 is an agreement table comparing illustrative variations of a human prediction method and prediction machine learning model.
[0114] FIG.86 are images of an illustrative variation of immunofluorescent stained ECMBs on a microfluidic chip.
[0115] FIG.87 is a receiver operating characteristic (ROC) curve of classifications by an illustrative variation of CRC prediction machine learning model.
[0116] FIG.88 is a binary classification matrix for classifications of eosin-stained ECMBs by an illustrative variation of an Advanced Adenoma (AA) prediction machine learning model.
[0117] FIG.89 is a table of model performance metrics at various thresholds for an illustrative variation of an AA prediction machine learning model trained to analyze eosin-stained samples.
[0118] FIG.90 is a binary classification matrix for classifications of antibody-stained ECMBs by an illustrative variation of an Advanced Adenoma (AA) prediction machine learning model.
[0119] FIG.91 is a table of model performance metrics at various thresholds for an illustrative variation of an AA prediction machine learning model trained to analyze antibody-stained samples.
[0120] FIG.92 depicts a flowchart representation of an illustrative variation of generating and selecting image parameters to generate a prediction machine learning model.Attorney Docket No.: AUMI-004 / 02WO 348385-2089
[0121] FIG.93 is a table of an illustrative variation of a training set comprising eosin-stained ECMBs associated with PDAC.
[0122] FIG.94 is a table of an illustrative variation of a test set comprising eosin-stained ECMBs associated with PDAC.
[0123] FIG.95 is a binary classification matrix for classifications of eosin-stained ECMBs by an illustrative variation of a PDAC prediction machine learning model.
[0124] FIG.96 is a table of an illustrative variation of a training set comprising eosin-stained ECMBs associated with early-stage PDAC.
[0125] FIG.97 is a table of an illustrative variation of a test set comprising eosin-stained ECMBs associated with early-stage PDAC.
[0126] FIG.98 is a binary classification matrix for classifications of eosin-stained ECMBs by an illustrative variation of an early-stage PDAC prediction machine learning model.
[0127] FIG.99 is a table of an illustrative variation of a training set comprising eosin-stained ECMBs associated with CRC.
[0128] FIG.100 is a table of an illustrative variation of a test set comprising eosin-stained ECMBs associated with CRC.
[0129] FIG.101 is a binary classification matrix for classifications of eosin-stained ECMBs by an illustrative variation of an CRC prediction machine learning model.
[0130] FIG.102 is a table of an illustrative variation of a training set and a test set comprising IF-stained ECMBs associated with PDAC.
[0131] FIG.103 is a binary classification matrix for classifications of IF-stained ECMBs by an illustrative variation of an PDAC prediction machine learning model.
[0132] FIG.104 is a table of an illustrative variation of a training set and a test set comprising IF-stained ECMBs associated with CRC.Attorney Docket No.: AUMI-004 / 02WO 348385-2089
[0133] FIG.105 is a binary classification matrix for classifications of IF-stained ECMBs by an illustrative variation of a CRC prediction machine learning model.
[0134] FIG.106 is a binary classification matrix for classifications of IF-stained ECMBs by an illustrative variation of an early-stage CRC prediction machine learning model.
[0135] FIG.107 are images of illustrative variations of gastric cancer ECMBs samples associated with a morphological feature.
[0136] FIG.108 are images of illustrative variations of gastric cancer ECMBs samples associated with another morphological feature.
[0137] FIG.109 are images of illustrative variations of esophageal cancer ECMBs samples associated with a morphological feature.
[0138] FIG.110 are images of illustrative variations of esophageal cancer ECMBs samples associated with another morphological feature.
[0139] FIG.111 are images of illustrative variations of liver cancer ECMBs samples associated with a morphological feature.
[0140] FIG.112 are images of illustrative variations of liver cancer ECMBs samples associated with another morphological feature.
[0141] FIG.113 are images of illustrative variations of liver cancer ECMBs samples associated with yet another morphological feature.
[0142] FIG.114 are images of illustrative variations of ECMBs corresponding control and cancerous ECMBs on a microfluidic chip. DETAILED DESCRIPTION
[0143] Described here are systems, devices, and methods for diagnosis, prognosis, and / or treatment planning of a subject based on extracellular matrix body (ECMB) (e.g., matricle) data corresponding to a biological fluid. For example, the systems, devices, and methods described herein may be useful for: non-invasive, high-throughput, scalable, and automated microfluidic processing and analysis of biological fluids; improving a signal strength for a pathology ofAttorney Docket No.: AUMI-004 / 02WO 348385-2089 interest by analyzing the ECMBs separated from the biological fluid; identifying and classifying one or more morphological features of ECMBs (e.g., abundance, distribution, shape, structure, color, texture, pattern, size, spatial localization, edges, edge properties) from a biological fluid; facilitating identification of disease ECMBs (e.g., early stage cancer, pre-cancer) and physiologically normal ECMBs with high sensitivity and specificity using one or more of a human (e.g., pathologist) and a prediction machine learning model; combining human and prediction machine learning model analysis of ECMB data; analysis of the ECMB data having one or more biomarkers using a prediction learning model to predict a cancer parameter (e.g., cancer type, cancer stage, pre-cancer); analysis of ECMB data using one or more histological stains with a prediction learning model to predict a cancer parameter (e.g., cancer type, cancer stage, pre-cancer); comparing a proteomic composition of ECMBs with a disease pathology; and facilitating disease monitoring by reducing the cost of sample collection and risk of complications associated therewith. Accordingly, morphological analysis of a biological fluid for cancer prediction (e.g., diagnosing cancer, assessing a cancer state, monitoring progression of cancer) as described herein may be performed faster, at lower cost, with higher throughput and frequency, and with higher stakeholder (e.g., patient, health care provider, caregiver, insurance provider) confidence than a conventional tissue biopsy. The subjects may or may not have developed cancer and may not be aware that they have cancer.
[0144] Conventional liquid biopsy diagnostic solutions are unreliable and suffer from poor accuracy of early-stage disease detection due to the low abundance of tumor markers in early- stage cancer and pre-cancer (e.g., pre-cancerous colon and rectal polyps, advanced adenoma, pre-cancer lesions), which may lead to false negatives. Conventional diagnostics, for example, may have insufficient sensitivity to detect early-stage or pre-cancer because the biomarker relied upon in conventional methods (e.g., ctDNA) are present at only low levels in early-stage cancer, making diagnosis challenging and often inaccurate. Additionally, conventional diagnostics based on DNA biomarkers may produce additional false positives by failing to distinguish between true cancer signals and non-malignant or age related mutations. Conversely, analysis of non- specific markers may result in false positives. For example, liquid biopsy assays lack the accuracy (e.g., sensitivity, specificity, positive and negative predictive values) to detect critical genetic alterations in earlier disease states due to the low concentrations of circulating tumor cell DNA (ctDNA), circulating tumor cells (CTCs), and / or extracellular vesicles (EVs) from a particular tumor. Moreover, conventional diagnostic solutions are prohibitively expensive (e.g.,Attorney Docket No.: AUMI-004 / 02WO 348385-2089 time, cost) for routine use in large-scale screening programs, thereby limiting their accessibility. These limitations in conventional techniques may delay diagnosis and administration of life- saving therapeutics.
[0145] Generally, the systems and devices described herein may include microfluidic systems configured for biological assays of ECMBs of a biological fluid. For example, a microfluidic system may include a microfluidic chip configured to receive a biological fluid (e.g., sample, blood) and separate ECMBs from the biological fluid. Separated ECMBs (e.g., nucleic acids, proteins, carbohydrates, lipids) may be exposed to one or more reagents and buffers for histochemical staining, immunohistochemistry staining, nucleic acid staining, and subsequent imaging and / or analysis. ECMB morphology (e.g., location of predetermined proteins and / or nucleic acids) may be spatially similar in non-cancerous states but distinct in disease states and / or pre-cancer. For example, non-cancerous homeostasis conditions exhibit a low variance in morphology (e.g., similar size, shape, distribution, interactions) while different disease states unique and distinct characteristics relative to non-cancerous ECMB morphology. In some variations, a proteomic and / or nucleic acid profile of ECMBs may be compared to biological controls for characterizing the ECMBs and their components for screening, diagnostic, prognostic capabilities, and therapeutic target identification.
[0146] In some variations, a system for predicting (e.g., diagnosing, monitoring, assessing) one or more of cancer and pre-cancer (e.g., pre-cancerous lesion) may receive image data corresponding to ECMBs separated from a biological fluid of a patient and generate corresponding ECMB data. In some variations, the ECMB data may comprise one or more biomarkers, morphological features of the ECMBs, or image parameters. A biomarker may correspond to a level of a substance found in the ECMBs including one or more proteins, polypeptides, lipid molecules, lipoparticles, carbohydrates, nucleic acid molecules, an expression level of a nucleic acid or protein biomarker, and the like.
[0147] In some variations, a prediction machine learning model may classify a sample as cancer or pre-cancer based on one or more morphological features or image parameters. A morphological feature or image parameter may generally be associated with one or more of an area, diameter, perimeter, shape descriptor (e.g., aspect ratio, circularity, solidity, or eccentricity), spatial location (e.g., location on chip, location within ECMB), color intensity, color distribution, texture features (including contrast, entropy, homogeneity, and uniformity),Attorney Docket No.: AUMI-004 / 02WO 348385-2089 edge property, structural features, intensity measurements, size distribution, abundance or count of structures, orientation, and biomechanical attributes (e.g., stiffness), and the like depending on prediction method and / or cancer parameter to be predicted.
[0148] For example, logistic regression analysis may be performed based on one or more of ECMB area, diameter, color, structure, texture and intensity. In some variations, the ECMB data may be analyzed using a prediction machine learning model to predict a cancer parameter such as a cancer type (e.g., colorectal cancer, lung cancer, pancreatic cancer, prostate cancer, breast cancer, esophageal cancer, liver cancer, gastric cancer, ovarian cancer, kidney cancer, melanoma), pre-cancer, and / or cancer stage. For example, the methods described herein include prediction machine learning models configured to provide sensitivity of greater than about 70%, greater than about 75%, greater than about 80%, greater than about 85%, greater than about 90%, greater than about 95%, and / or greater than about 99% for colorectal cancer from non- cancerous controls. The prediction machine learning model may be trained using a training set of stained (e.g., eosin, antibody, nucleic acid stain, or any combination thereof) ECMB data based on one or more of abundance, morphology, and spatial localization (e.g., location of an ECMB on the chip, intra-ECMB locations) of the separated ECMBs. In some variations, ECMB data may be based on one or more image parameters. The one or more image parameters may be associated with the abundance, morphology (e.g., morphological features), and / or spatial localization of the separated ECMBs. Additionally or alternatively, a pathologist or trained observer may score the morphological features of the ECMB data to predict a cancer parameter.
[0149] Sensitivity corresponds to an ability of a prediction machine learning model to predict that an ECMB sample is positive (e.g., cancerous) by dividing the number of true positives by the number of true positives plus the number of false negatives. For example, a highly sensitive model provides fewer false negative results, and thus fewer cases of disease may be missed. Specificity corresponds to an ability of a prediction machine learning model to predict that an ECMB sample is negative (e.g., non-cancerous) by dividing the number of true negatives by the number of true negatives plus the number of false positives. For example, specificity corresponds to the percentage of true negatives out of all subjects who do not have a disease or condition. Positive predictive value (PPV) corresponds to a number of true positives out of all positive predictions (e.g., true positives and false positives). Negative predictive value (NPV)Attorney Docket No.: AUMI-004 / 02WO 348385-2089 corresponds to a number of true negatives out of all of negative predictions (e.g., true negatives and false negatives).
[0150] A biological fluid may comprise one or more of human or other animal bodily fluids, tissue, cells, whole blood, blood plasma, blood serum, cerebrospinal fluid, intrathecal fluid, urine, saliva, sweat, tears, synovial fluid, pleural fluid, gastric fluid, peritoneal fluid, breast milk, nipple aspirate, semen, amniotic fluid, vitreous, aqueous humor, lymph, bile, cerumen, chyle, chyme, endolymph, perilymph, exudates, feces, ejaculate, gastric acid, gastric juice, mucus, pericardial fluid, pus, rheum, sebum, serous fluid, smegma, sputum, synovial fluid, vaginal secretion, menstrual effluent, vomit, tumors, conditioned cell culture media, carriers, reagents, solutions, binding moieties, and the like. As contemplated herein, the systems, methods, and devices of the disclosure are useful for diagnosis, prognosis, and the like from human and other animal subjects, including but not limited to one or more of non-human primates, companion animals (e.g., cats, dogs, or other pets) and livestock (e.g., cattle, sheep, porcine, horses, goats, and the like).
[0151] A buffer may comprise one or more of MES, HCL acid buffer, Acid Phthalate buffer, Alkaline borate buffer, Acetate buffer, Acetic ammonia buffer, acetone buffer, ammonia buffer, barbitone buffer, buffered copper sulfate solution, glycerin solution, glycine buffer solution, palladium chloride buffer solution, citric acid Na2HPO4, citric acid Sodium Citrate Buffer Preparation, Sodium Acetate Acetic Acid Buffer Preparation, Na2HPO4-NaH2PO4, Imidazole (glyoxaline), Sodium Carbonate, TBE, TAE, BIS-TRIS, Bis-Tris propane, Phosphate buffer, formic acid, Pyridine and conjugate acid, Ammonia and conjugate acid, Methylamine and conjugate acid, ADA, ACES, PIPES, MOPSO, BES, MOPS, TES, HEPES, DIPSO, MOBS, TAPSO, Tris, Trizma, HEPPSO, POPSO, TEA, EPPS, Tricine, Gly-Gly, Bicine, HEPBS, TAPS, AMPD, TABS, AMPSO, CHES, CAPSO, AMP, CAPS, CABS, and the like.
[0152] Some microfluidic chips and systems suitable for use in the systems here are described in International Patent Application No. PCT / IB2024 / 062532, filed December 11, 2024, and titled “SYSTEMS, DEVICES, AND METHODS FOR MICROFLUIDIC FLUID ANALYSIS,” U.S. Patent Application No.63 / 608,790, filed on December 11, 2023, and titled “SYSTEMS, DEVICES, AND METHODS FOR MICROFLUIDIC FLUID ANALYSIS,” and International Patent Application No. PCT / US2021 / 023827, filed March 24, 2021, and titled “DEVICE AND METHODS FOR ISOLATING EXTRACELLULAR MATRIX BODIES,” International PatentAttorney Docket No.: AUMI-004 / 02WO 348385-2089 Application No. PCT / US2019 / 052310, filed September 21, 2019, and titled “COMPOSITIONS AND METHODS FOR GLAUCOMA,” each of which is hereby incorporated by reference in its entirety. I. Methods
[0153] Described here are methods of diagnosis, prognosis, and / or treatment of a subject. The methods described herein may be useful in providing high throughput analysis with high sensitivity and specificity of a cancer parameter (e.g., cancer type, cancer stage, pre-cancer) based on ECMBs separated from a biological fluid such as blood to assist diagnosis and / or development of a treatment plan for a subject, and to monitor the effectiveness of a treatment plan. Furthermore, the methods may compare the proteome of a biological fluid to both the ECMB fraction isolated from the biological fluid and the eluted fraction of the biological fluid. For example, a predetermined fraction of the biological fluid or the ECMBs may be analyzed to determine one or more characteristics of a biomarker, morphological feature, and / or image parameter associated with a disease, which may be useful in diagnosis and treatment selection. The biomarkers, morphological features, and / or image parameters may be used not only for cancer screening and prediction, but also to support cancer prognosis, facilitate treatment selection, monitor minimal residual disease, subtype patients for clinical trials, and monitor both the safety and therapeutic signals in subjects after treatment has begun, and during the course of treatment. For example, the methods described herein may guide personalized therapies for treatment of a subject using minimally invasive or noninvasive biological fluid sampling.
[0154] FIG.1 is flowchart that generally describes a method of predicting one or more of cancer and pre-cancer from a biological fluid 100 using any of the systems and devices described herein. The method 100 may include transferring a biological fluid to a microfluidic chip 102. The microfluidic chip (e.g., chip 316) may be configured to process the biological fluid to separate ECMBs from the biological fluid in a manner that maintains their composition and properties (e.g., morphology) to facilitate their use as biomarker(s) for disease diagnosis and / or monitoring of chemical or biological processes. For example, the systems described herein including system 1200, 1300 depicted in respective FIGS.12, 13A, and 13B may provide high throughput separation of extracellular matrix bodies from a biological fluid. The microfluidic chips described herein include microfluidic chip 1416 depicted in FIG.14. TheAttorney Docket No.: AUMI-004 / 02WO 348385-2089 biological fluid may comprise one or more of whole blood, blood plasma, and blood serum. For example, a biological fluid may be loaded into a microfluidic chip 1216 using a robot 1214.
[0155] In some variations, the ECMBs in the microfluidic chip may be processed 104. In some variations, the processing the ECMBs (e.g., matricles) in the microfluidic chip may include applying one or more of a histochemical stain, an immunohistochemical (IHC) stain, immunofluorescence (IF) stain, a multiplex IHC or IF stain, multi-spectral imaging, a protein stain, a glycomic stain, a nucleic acid stain, chemical fixation, wash solution, buffer solution, or mounting media, and a protease inhibitor. For example, human plasma may be stained with histological reagents including hematoxylin and eosin, 2',7'-dichlorofluorescin diacetate, 7- amino-actinomycin D, Acid Fast Bacteria, Acid Fuchsin (Acid Violet 19), Acid Orcein, Acid Schiff's, Acridine Orange (Basic Orange 14), Acridine Yellow (Basic Yellow K), Actin and Tubulin Dyes, Alcian Blue 8GX (Ingrain Blue 1), Alcian Yellow (Red 83), Aldehyde fuchsin, Alizarin Red S (Mordant Red 3), Alkaline phosphatase, Aniline Blue (H20 Sol) (Sol blue 3 M or 2R, water blue) (Acid Blue 22), Auramine 0 (Basic Yellow 2), Azan Stain, Azophloxine (Acid Red 1), Azure A (McNeal), Biebrich Scarlet (Acid Red 66), Bielshowsky stain, Bismark Brown Y (VesuvianBrown) (Basic Brown 1), Bromodeoxyuridine (5-Bromo-2′-deoxyuridine), Brown & Brenn, Calcein-AM, Calcofluor-white, Carminic acid (Natural Red 4), CellTracker Dyes: (e.g., CellTracker Green CMFDA, CellTracker Orange CMTMR), Chrome Azurol S, Chromotrope 2R (Acid Red 29), Colloidal Iron Stain, Congo Red (Direct Red 28), Coomassie Brilliant Blue (Acid Blue 83), Cresyl Fast Violet, Crystal (Basic Violet 3), Crystal Ponceau 6R (brilliant Crystal Scarlet 6R, Ponceau 6R) (Acid Red 44), Crystal Violet, DAPI (4',6-diamidino- 2-phenylindole), Diaphonization, DiB, DiD, DiI, Dil, DiO, Elastin Verhoeff Vangieson (VVG), Eosin, Eosin Bluish (Eosin B, Erythrosin B) (Acid Red 51), Eosin Yellowish (H20 & alcohol Sol, Eosin Y) (Acid Red 87), ER-Tracker Dyes, Ethidium Bromide, Ethyl Green, Fast Garner GBC salt (Azoic Diazo component 4), Fast Green FCF (Food Green 3), Fast Red B salt (Azoic Diazo component 5), Fast Red TR salt (Azoic Diazo component 11), Feulgen Stain, Fites Acid Fast, Fluorescein (Acid Yellow 73), Fluorescein isothiocyanate, Fluorescently labeled DNA probes complementary to DNA-Conjugated antibodies, Fluorescently-conjugated antibodies for dsDNA, ssDNA and RNA, FM1-43, FM4-64, Fontana Masson – melanin, Fuchsin Acid (Acid Violet 19), Fuchsin basic (Basic Violet 14), Fuchsin Carbol, Fuchsin new (Basic Violet 2), Gallocyanine (Mordant Blue 10), Gentian Violet, Giemsa Stain, Gimenez / Pierce Vanderkamp, Golgi-Tracker Dyes, Gomori Trichrome, Gram Stain, Hall's Bilirubin, Hematein, HematoxylinAttorney Docket No.: AUMI-004 / 02WO 348385-2089 (Natural Black 1), Hemosiderin (Fe), Hoechst 33258, Hoechst 33342, Indigo carmine (Food Blue 1), IRON Gomori's, Janus Green B, Jaswant Singh-Bhattacharji, Kinyoun's Acid Fast, Lactophenol cotton blue, Light Green SF (Acid Green 5), Lipofuscin – AFIP, Liu's stain, Luna stain, Luxol Fast Blue (Solvent Blue 38), LysoTracker Dyes, Machiavello, Malachite Green (Basic Green 4), Mallory trichrome, Mallory's Phosphotungstic Acid Hematoxylin, Mallory's Triple Stain, Martius Yellow (Acid Yellow 24), Masson's Trichrome, Mayer's Mucicarmine, Melanin Bleach, Melanin stain (Fontana-Masson), Metanil Yellow (Acid Yellow 36), Methyl Blue (Acid Blue 93), Methyl Green (Basic Blue 20), Methyl Violet 2B (Basic Violet 1), Methylene Blue (Basic Blue 9), Methyl Green Pyronin, MitoTracker Dyes, Movat's Pentachrome, Mucicarmine, Neuro-DiO, Neutral Red (Basic Red 5), Nigrosin, Nile Blue Oxazone, Nile Blue sulfate (Basic Blue 12), Nissl, Oil Blue 35, Oil Red O (Solvent Red 27), Orange G, Orcein (Natural Red 28), Osmium tetroxide, Pacific Blue, Papanicolaou Stain, PAS – basement membrane, PAS / Diastase, Patent Blue (Acid Blue 1), Periodic acid meth silver stain, Periodic acid schiff, Periodic acid schiff with dig, Perl's Iron, Phalloidin Conjugates, Phloxin (Acid Red 92), Phosphine (Basic Orange 15), Phosphotungstic Acid Hem, Phyloxin, Picric acid, Picrosirius red, Ponceau 2R (Ponceau de xylidene) (Acid Red 26), Propidium Iodide, Prussian Blue, Pyranine, Pyronin Y (Pyronin G), QDs (Quantum dots in imaging in cells including epifluorescence, confocal and multiphoton), Quinoline Yellow SS (Solvent Yellow 33), Red 2G, Red Oil 3, Reticulin stain, Rhodamine, Rhodamine 123, Rhodamine 6G, Rhodamine B (Basic Violet 10), Rhodanine – copper, RiboGreen, Romanowsky stains, Rose Bengal (Acid Red 94), Ruthenium Red, SAFO Safranin & O, Safranin O (Basic Red 2), Scarlet R (Solvent Red 24), Silver Nitrate, Silver Stains, Sirius Red (Direct Red 80), SiR-Tubulin, Solochrome cyanine RS (Eriochrome cyanine R) (Mordant Blue 3), Sudan Black (Solvent Black 3), Sudan III, Sudan IV, Sudan Red 7B (Solvent Red 19), SYBR Gold (cyanine dye), SYBR Green I (cyanine dye), SYBR Safe (cyanine dye), SYTOX Green, Tartrazine (Food Yellow 4), Texas Red (sulforhodamine 101 acid chloride), Thioflavine T (Basic Yellow 1), Thionin S, Toluidine Blue (Basic Blue 17), Trichrome – Masson blue, Tubulin Tracker Green, Uzman, Van Gieson's Stain, Verhoeff-Van Gieson, Victoria Blue B (Basic Blue 7), Von Kossa – calcium, Warthin-Starry, Wayson stain, Weigert's Resorcin-Fuchsin, and Ziehl-Neelson Acid Fast.
[0156] Examples of suitable fluorescent cell dyes include DAPI (4',6-diamidino-2- phenylindole), Hoechst 33342 and Hoechst 33258, Propidium Iodide (PI), SYTOX Green, Acridine Orange, DiI, DiO, DiB, Neuro-DiO, DiD, CFSE, FM1-43, FM4-64, Calcein-AM,Attorney Docket No.: AUMI-004 / 02WO 348385-2089 CellTracker dyes (e.g., CellTracker Green CMFDA, CellTracker Orange CMTMR), MitoTracker dyes, LysoTracker dyes, ER-Tracker dyes, Golgi-Tracker dyes, Actin and Tubulin dyes, Phalloidin conjugates, Tubulin Tracker Green: SiR-Tubulin, DCFDA (2',7'- dichlorofluorescin diacetate), fluorescently labeled DNA probes complementary to DNA- Conjugated antibodies, fluorescently-conjugated antibodies for dsDNA, ssDNA and RNA, and QDs (e.g., Quantum dots in imaging in cells including epifluorescence, confocal and multiphoton).
[0157] In some variations, the stains and / or dyes (e.g., histochemical stain, immunofluorescence (IF) stain, protein stain, etc.) applied to a sample processed on a microfluidic chip may be selected based on predetermined proteins. For example, the predetermined proteins may be associated with cancerous samples. Both the predetermined proteins and selected stains and / or dyes may vary depending on cancer type (e.g., colorectal cancer, breast cancer, lung cancer, prostate cancer, esophageal cancer, liver cancer, gastric cancer, ovarian cancer, kidney cancer, melanoma). In some variations, the predetermined proteins may be associated with one or more Gene Ontology (GO) terms relevant to the biological processes, cellular components, and / or molecular components that vary between cancerous and non-cancerous samples and are thus informative with respect to a cancer or non- cancerous classification. In some variations, the predetermined proteins may be associated with one or more Gene Ontology (GO) terms corresponding to the cellular components (e.g., ECMBs) processed on the microfluidic chip. For example, the stains and / or dyes applied to a sample may be selected based on predetermined proteins associated with GO:0070062 relating to extracellular exosomes, GO: 0005615 relating to extracellular space, and / or GO:0005576 relating to extracellular regions.
[0158] In some variations, the ECMBs in the microfluidic chip may be processed absent staining. The biological fluid may be processed using one or more of microfluidic separation, chemical fixation, physical fixation, affinity chromatography, sedimentation, centrifugation, differential centrifugation, density gradient centrifugation, mesh filtration, diafiltration, tangential flow filtration, membrane filtration, elutriation, affinity-based capture, precipitation, immuno-affinity capture, tag-based affinity capture, protein A / G / L affinity capture, enzyme or receptor-ligand affinity capture, lectin-based affinity capture, carbohydrate and sugar-based affinity capture, capture by carbohydrate-binding modules, synthetic glycopolymer affinityAttorney Docket No.: AUMI-004 / 02WO 348385-2089 capture, hybridization-based capture, aptamer-based affinity capture, affinity capture based on peptide nucleic acid probes, affinity capture based on protein-nucleic acid interactions, magnetic bead capture, ultrasonic capture, size exclusion chromatography, ion exchange chromatography, hydrophobic interaction chromatography, electrophoresis, dialysis, flow cytometry, field-flow fractionation, AC electrokinetics, and embedding (e.g., embedding in one or more of paraffin, wax, freezing media, celloidin, nitrocellulose, agar, agarose, gelatin, and the like). In some variations, one or more of the above processing techniques may be combined with any one or more of the staining processes described herein to process ECMBs.
[0159] In some variations, the microfluidic chip may be configured to separate ECMBs from the biological fluid, and thereafter the microfluidic chip and separated ECMBs may be imaged using transmitted light microscopy without staining. For example, in some variations, the chip may be imaged using single channel transmitted light microscopy and / or multichannel transmitted light microscopy (e.g., 8 channel brightfield microscopy). In some variations, the chip may be imaged using one or more of brightfield, darkfield or scattering contrast, Rheinberg illumination, oblique illumination, gradient contrast, phase contrast, differential interference contrast, Nomarski, Shear, polarized light or birefringence contrast, Hoffman Modulation Contrast or relief contrast, and diffraction phase microscopy. In some variations, contrast may be generated in a chip image with computational methods using one or more of quantitative phase imaging (QPI), including transport of intensity equation (TIE), spatial light interference microscopy (SLIM), diffraction phase microscopy (DPM), digital holographic microscopy (DHM), and differential phase contrast (DPC). In some variations, contrast may be generated with computational illumination and synthetic aperture methods using one or more of Fourier ptychography (FP), structured illumination microscopy (SIM), computational reconstruction for SR and contrast. In other variations, transmitted light microscopy may be used to image microfluidic chips with one or more stains and / or dyes, as described herein, applied to the ECMBs on the microfluidic chip.
[0160] For example, FIGS.2A-2E are images of processed ECMBs (e.g., matricles) on respective microfluidic chips including ECMBs corresponding to respective control plasma ECMBs and cancer plasma ECMBs. For example, the images 200, 202, 204, 206, 210, 212, 214, 216,, 224, 226, 220, 222, 230, 232, 234, 236, 240, 242, 244, 246 in FIGS.2A-2E correspond to biological fluids (e.g., blood plasma samples) processed on respective microfluidic chipsAttorney Docket No.: AUMI-004 / 02WO 348385-2089 including pillars 201 of decreasing inter-pillar distance (e.g., 100 μm gap region, 50 μm gap region, 25 μm gap region, 15 μm gap region, 4 μm gap region). For example, the biological fluid flows into the 100 μm gap region and toward the 4 μm gap region such that insoluble material is trapped in the spaces between the pillars. The ECMBs are eosin-stained (e.g., Eosin Y). In images 200, 202, 204, 206 of FIG.2A, plasma ECMBs 203, 205 from non-cancerous control plasma localize predominantly to the 4 μm gap region of the microfluidic chip. Non-cancerous control ECMBs may generally be small and do not organize as large clusters. In images 210, 212, 214, 216 of FIG.2B, pancreatic cancer ECMBs 211, 213, 215 from a plasma sample localize substantially across the entire microfluidic chip and may have a large sheet-like shape. In images 220, 222, 224, 226 of FIG.2C, colorectal cancer ECMBs 211, 213, 215 from a plasma sample localize substantially across the entire microfluidic chip with a higher concentration shown in the large gap region of magnified image 222. In images 230, 232, 234, 236 of FIG.2D, prostate cancer ECMBs 231, 233, 235 from a plasma sample localize substantially across the entire microfluidic chip with a higher concentration shown in the large gap region of magnified image 232. Some prostate cancer ECMBs have a long and fibrous shape and may be organized as large clusters of discreet puncta. In images 240, 242 of FIG.2E, lung cancer ECMBs 241, 243, 245 from a plasma sample localize substantially across the entire microfluidic chip with a slightly higher concentration in the small gap region of magnified image 242. Some lung cancer ECMBs may be fibrous and overlaid with adjacent puncta.
[0161] As another example including additional cancer types, FIG.114 are transmitted light microscopy images 11402-11418 and 11440 of processed ECMBs on respective microfluidic chips. The ECMBs of images 11402-11418 and 11440 were imaged absent staining. Image 11440 of FIG.114 corresponds to biological fluids (e.g., blood plasma samples) of a healthy control sample processed on respective microfluidic chips including pillars. Images 11402, 11404, 11406, 11408, 11410, 11412, 11414, 11416, and 11418 of FIG.114 corresponded to biological fluid associated with breast cancer, ovarian cancer, esophageal cancer, gastric cancer, liver cancer, pancreatic cancer, lung cancer, kidney cancer, and melanoma respectively, each processed on respective microfluidic chips. All samples shown in FIG.114 were acquired prior to chemotherapeutic intervention from a cohort of 45-80 year-old patients. Samples from patients with a history of malignancy were excluded. Plasmas were collected in K2 EDTA tubes and subject to centrifugation from 1000g to 3000g. To isolate extracellular matrix bodies (ECMBs), plasma was perfused through the microfluidic chip, followed by washing with PBS.Attorney Docket No.: AUMI-004 / 02WO 348385-2089 Images were acquired using brightfield microscopy at 10x using Axioscan 7 and an Axiocam 705 color camera. As described in more detail herein, processing and separating the cancerous ECMBs of images 11402-11418 may reveal clear differences in morphological features when compared with the morphological features of the control ECMBs of image 11440 or with cancers of differing cancer types.
[0162] In some variations, one or more predetermined antibodies (e.g., antibodies specific for extracellular matrix materials, cancer-related markers, extracellular vesicle markers), washes, and reagents may be applied to the microfluidic chip to enable IHC staining. Applying such treatments to ECMBs on a microfluidic chip may improve cancer diagnosis compared to traditional gene sequencing by providing additional information about the spatial localization and morphology of the cellular material (e.g., ECMBs) of a given sample. Unlike proteomic methods that rely on homogenized tissue, the methods described herein preserve spatial localization of where each protein is expressed within an ECMB. This combination of protein expression abundance, as indicated by signal intensity, and precise spatial localization may provide a more informative dataset for improved prediction and diagnostic accuracy.
[0163] For example, a plurality of antibodies (e.g., antibodies specific for extracellular matrix, cancer-related markers, extracellular vesicle markers) may be applied to the microfluidic chip for IHC or IF multispectral imaging. For example, FIG.86 shows images of processed ECMBs (e.g., matricles) stained with immunofluorescent (IF) stains on a microfluidic chip. FIG.86 contains images 8602, 8604, 8606, 8608, and 8610 corresponding to biological fluids (e.g., blood plasma samples) processed on a microfluidic chip including a pillar 8620. The ECMBs 8632, 8634, 8636, 8638, and 8640 correspond to stage 1 breast cancer plasma. Images 8602, 8604, 8606 and 8608 of FIG.86 show cancer associated ECMBs 8632, 8634, 8636, and 8638 stained with a single IF stain. For example, image 8604 shows ECMBs 8634 stained with an antibody stain and image 8610 shows ECMBs 8640 stained with a DAPI stain. In some variations, ECMBs on a microfluidic chip may be processed with a plurality of IF stains such as a preconfigured multiplex immunofluorescence stain (mIF) for use with the cancer parameter prediction methods described herein. Image 8602 of FIG.86 shows cancerous ECMBs 8632 on a microfluidic chip stained with a multiplex immunofluorescent stain.
[0164] In some variations, successive staining of an isolate fraction on the microfluidic chip may help remove the material within the microfluidic chip. For example, antibody staining of aAttorney Docket No.: AUMI-004 / 02WO 348385-2089 sample may generate a first spatial arrangement on the microfluidic chip. If the microfluidic chip is subsequently stained (e.g., with a different set of reagents), a second spatial arrangement different from the first spatial arrangement may be generated due to movement of the sample within the microfluidic chip. Thus, a useful comparison between the first and second spatial arrangements may be challenging. Accordingly, chemical fixation (e.g., crosslinking) of biological fluid (e.g., ECMBS) on the chip may be applied to immobilize the biological fluid in the microfluidic chip. For example, the microfluidic chip may receive a carbodiimide fixative and aldehyde fixative configured to crosslink the biological fluid to immobilize the biological fluid on the microfluidic chip. In some variations, one or more of the following chemical fixatives may be used: cross-linking fixatives such as formaldehyde, glutaraldehyde, carbodiimide, diimidoesters; denaturing fixatives such as alcohol, ethanol, methanol, acetone, chloroform, glacial acetic acid, Carnoy’s fixative, Clarke’s fluid, methyl Carnoy, potassium permanganate; metallic fixatives such as mercuric chloride-base fixatives, osmium tetroxide, potassium dichromate-based fixatives, zinc salts, uranyl acetate; and picric acid-based fixatives such as Bouin’s solution, Hollande’s Modified Bouin’s Solution, Gendre’s Fluid, Rossman’s Fluid, Zamboni’s Fluid, Duboscq-Brazil Fluid. In some variations, biological fluid can be immobilized using one or more physical fixation methods, such as cryofixation, high-pressure freezing (HPF), lyophilization, freeze-substitution, heat fixation, direct flame heat fixation, microwave fixation, microwave-assisted chemical fixation, air drying, and vapor fixation.
[0165] In some variations, processing the biological fluid may include using one or more of microfluidic separation, affinity chromatography, centrifugation, differential centrifugation, density gradient centrifugation, mesh filtration, diafiltration, tangential flow filtration, membrane filtration, immuno-affinity capture, magnetic bead capture, size exclusion chromatography, electrophoresis, field-flow fractionation, and AC electrokinetics.
[0166] In some variations, one or more biomarkers in one or more of the biological fluids and the ECMBs may be analyzed by one or more of microscopy, microfluidic device, mass spectrometry, microarray, nucleic acid amplification, hybridization, proteomic profiling, fluorescence hybridization, immunohistochemistry, nucleic acid analysis or sequencing, next generation sequencing, flow cytometry, chromatography, electrophoresis, immunostaining, fluorescence assay, fluorescent in situ hybridization (FISH), chelate complexation, quantitative HPLC, spectrophotometry, colorimetric assay, chemiluminescence assay, immunofluorescenceAttorney Docket No.: AUMI-004 / 02WO 348385-2089 assay, light scattering, infrared, UV-VIS, antibody array, Western blot, immunoassay, immunoprecipitation, ELISA, LC-MS, LC-MRM, radioimmunoassay, 2D gel mass spectrometry, LC-MS / MS, RT-PCR, and quantitative PCR. In some variations, the biological fluid may be analyzed as one or more of a total biological fluid fraction, an isolate fraction, and an eluate fraction.
[0167] In some variations, the separated ECMBs may comprise an isolate fraction. For example, a microfluidic chip 1216, 1416 may be washed to remove unincorporated antibodies, thereby separating the ECMBs from the biological fluid. The ECMBs in the microfluidic chip 1216, 1416 may be fixed and labeled with DAPI prior to a final wash. In some variations, the ECMBs on the microfluidic chip 1216, 1416 may undergo histochemical staining using Eosin Y, which is a general-purpose histopathological dye used to identify proteinaceous material. Staining may follow SOP QC-022, summarized as follows: Eosin Y aqueous solution is prepared by adding glacial acetic acid (e.g., about 20 µL) dropwise to Eosin Y (e.g., 0.5 w / v %, about 10 mL). The unfiltered Eosin solution (e.g., about 100 µL) is diluted with about 300 µL of MilliQ water to produce Eosin Y (e.g., 0.125 w / v %, 400 µL), which is filtered with a 0.22 µm filter prior to use. After washing the microfluidic chips 1216, 1416 with about 20 µl of MilliQ water, about 20 µl Eosin (e.g., 0.12 w / v%) may flow through the microfluidic chip 1216, 1416 and be incubated (e.g., stained) for about 3 minutes. Thereafter, the microfluidic chip 1216, 1416 may be washed with about 20 µl MilliQ water and processed for imaging.
[0168] In some variations, proteomic analysis of one or more of an unfixed or fixed isolate fraction, an eluate fraction, and an unfractionated biological fluid may be performed. For example, the eluate fraction (e.g., about 400 µL) of a biological fluid may be collected from an outlet of a microfluidic chip 1216, 1416. An isolate fraction including the separated ECMBs in the microfluidic chip 1216, 1416 may be suspended in about 1% SDS / PBS and then flushed out of the microfluidic chip 1216, 1416 to obtain an isolate fraction (e.g., about 500 µL). One or more of the isolate fraction, the eluate fraction, and the unfractionated biological fluid may be centrifuged at about 2,000 g for about 10 minutes, and then centrifuged at about 12,000 g for about 10 minutes at about 4ºC. The supernatant may be discarded and the pellet may be resuspended in about 1% SDS and pelleted again at about 25,000 g. The supernatant and pellet may be subjected to electrophoresis into NuPAGE 10% Bis-Tris gel. Peptides may be extracted from each gel band.Attorney Docket No.: AUMI-004 / 02WO 348385-2089
[0169] Samples may be analyzed by nanoscale liquid chromatography coupled to tandem mass spectrometry (nano LC-MS / MS) and an in-line analytical column. The number of total peptides (e.g., spectral count) in each sample from the human plasma may be determined, and used to normalize each sample. Using the normalized number of peptides in all samples, the protein expression profiles of a disease sample and a control sample, as well as the protein expression profiles of disease isolates and control isolates may be determined. For example, the fold change of total proteins in the chip eluates, and chip isolates in disease samples compared to control samples may be determined. The fold change of proteins in both disease and control samples enriched by the microfluidic chip 1216, 1416 (e.g., enriched in the eluate fraction and enriched in the isolate fraction) may also be determined. In some variations, differentially expressed proteins and total proteins in each fraction may be analyzed for enrichment (e.g., analyzed using: gene ontology (GO) enrichment analysis, pathway enrichment analysis, domain enrichment analysis, gene set enrichment analysis (GSEA), functional class scoring (FCS), network topology-based analysis (NTA), protein-protein interaction (PPI) network analysis, post- translational modification analysis, and combinations thereof).
[0170] In some variations, the separated ECMBs may undergo multiplex immunohistochemical staining by thawing biological fluid aliquots and incubating them with a mix of a plurality of antibodies (e.g., about 2, 3, 4, 5, 6, or more) corresponding to one or more biomarkers. For example, the biomarkers may be selected based on proteomic analysis of cancer ECMBs and non-cancerous ECMBs. In some variations, the antibodies may be directly conjugated to different fluorescent dyes to facilitate multiplex immunohistological (e.g., IHC, immunofluorescent) imaging. In some variations, an antibody labeled biological fluid (e.g., plasma) may be loaded onto a microfluidic chip 1216, 1416 for immobilizing immunostained material for subsequent morphological microscopic analysis.
[0171] In some variations, successive staining of processed ECMBs (e.g., matricles) on a microfluidic chip may label different components of the ECMBs for use in the prediction of a caner parameter. Either of a preceding or a subsequent staining of a staining sequence may include a combination or mixture of stains. For example, a sample may be first processed with a histological stain (e.g., Eosin) and then processed with a nucleic acid stain (e.g., DAPI). The subsequent staining may include a mixture of immunofluorescence (IF) antibodies, as described in more detail herein. One or more of the sequence of stains, the preceding stain, and theAttorney Docket No.: AUMI-004 / 02WO 348385-2089 subsequent stain may be associated with a predetermined cancer parameter prediction. For example, a predetermined IF antibody stain mixture may be associated with a prediction machine learning model corresponding a predetermined caner type (e.g., pancreatic cancer) to produce high sensitivity method of predicting a cancer parameter.
[0172] In some variations, image data corresponding to the ECMBs separated from a biological fluid of a subject may be received 106. For example, separated ECMBs on the microfluidic chip 1216, 1416 may be imaged based on brightfield imaging techniques to obtain spatial morphological information corresponding to the distribution of ECMBs on the chip. Additionally or alternatively, separated ECMBs on the microfluidic chip 1216, 1416 may be imaged based on fluorescence imaging of antibodies to obtain quantitative spatial morphological information on the distribution of predetermined biomarkers. For capturing the image data, predetermined exposure ranges may be employed when imaging the separated ECMBs on the microfluidic chip. For example, brightfield or color images may be taken with exposure times (e.g., shutter speeds) of about 0.5 microseconds to about 2 milliseconds, including all subranges and values therein. The exposure may be adjusted to achieve optimal illumination without saturation, thereby preserving true color representation and morphological detail. For fluorescent images, a signal to noise ratio may be weaker and background noise minimization may be describable. To reduce background noise, exposure times of fluorescent images may be longer compared to color images, for example from about 1 to 500 milliseconds, including all subranges and values therein, to ensure sufficient photon collection from fluorophores, minimize photobleaching, maximize the signal-to-noise ratio.
[0173] In some variations, one or more regions of interest may be identified in the image data. For example, a cancer parameter may be predicted based on one or more regions of interest including less than all ECMBs in the image data. The region of interest may include one or more pillars of a microfluidic chip. In some variations, the region of interest may be sized such that it includes a portion of an ECMB and does not include a pillar. In some variations, the region of interest may be identified by one or more of a trained human (e.g., pathologist, scientist), image analysis, and a prediction machine learning algorithm. For example, a region of interest may be manually identified by a human (e.g., morphology trained pathologist) and / or analyzed by a prediction machine learning model. In some variations, the region of interest may include one or more morphological features. In some variations, the manually identified regions of interest mayAttorney Docket No.: AUMI-004 / 02WO 348385-2089 be compared to the prediction machine learning model analyzed regions of interest, and may further increase the confidence of a cancer parameter prediction. For example, a variance between the identified regions of interest below a predetermined threshold may correspond to an increase in the confidence of a cancer parameter prediction. Conversely, a variance between the identified regions of interest above a predetermined threshold may correspond to a decrease in the confidence of a cancer parameter prediction, which may lead to additional analysis and / or testing.
[0174] In some variations, ECMB data may be generated based on the image data 108. In some variations, the ECMB data may include at least one morphological feature. In some variations, a morphological feature may correspond to one or more of a geometry, a size, a location, an abundance, a color, a texture, an edge, and a stiffness of the ECMB. For example, a morphological feature corresponding to the color of a processed and stained ECMB may distinguish an ECMB sample as cancerous or non-cancerous and / or indicate a stage of cancer. In some variations, expression of a morphological feature may vary between an ECMB associated with cancer and an ECMB associated with non-cancer. A morphological feature may be analyzed by one or more of a pathologist and prediction machine learning model when predicting a cancer parameter. In some variations, the ECMB data may include at least one image parameter. An image parameter or plurality of image parameters may be generated as described herein with steps 9202-9208 of method 9200 described with respect to FIG.92.
[0175] In some variations, ECMB data may be generated from a region of interest identified in the received image data. The region of interest may include one or more morphological features. In some variations, at least one morphological feature of the ECMBs may be segmented by inputting the ECMB data into the prediction machine learning model. The prediction machine learning models are described in more detail herein and with respect to FIGS.3A, 3B, and 92. In some variations, the ECMB data may include one or more image parameters, as described herein in more detail. In some variations, the ECMB data may comprise one or more biomarkers selected from the group consisting of Fibronectin, Tetranectin, Thrombospondin, Galectin-3 Binding Protein (3BP), Talin, Rab27B, Zyxin, DAPI, CD5L, Afamin, Carbonic Anhydrase (CA1), INF2, Clusterin, Victronectin, Gelsolin, S100A9, CD5L, and combinations thereof.
[0176] In some variations, the ECMB data may comprise one or more of a flake, a punctum, a decorator, and combinations thereof. For example, the flake may comprise an eosinophilicAttorney Docket No.: AUMI-004 / 02WO 348385-2089 structure having an area of between about 5 μm2and about 450 μm2. The puncta may comprise a circular structure having an area less than about 15 μm2. The decorator may comprise a circular structure having an area less than about 7 μm2.
[0177] In some variations, a cancer parameter may be predicted based on at least the ECMB data 110. In some variations, the cancer parameter may be predicted based on one or more of the received image data and clinical data in addition to the ECMB data. For example, clinical data may include a history of present illness (HPI), patient age, gender, medical history, surgical history, medications, presenting symptoms, diagnosis, laboratory results, prior imaging studies, relevant medications, clinical notes, family history, and risk factors. In some variations, clinical data may be preprocessed and / or encoded before being analyzed by a prediction machine learning model to predict a cancer parameter.
[0178] For example, the ECMB data may be analyzed using a prediction machine learning model, as described in more detail herein. In some variations, the prediction machine learning model may comprise one or more of architectures including but not limited to ResNet, Inception, boosting algorithms, bootstrap aggregation, random forests, decision trees, Vision Transformers, convolutional neural networks (CNNs), or other classical and deep learning models suitable for image analysis and classification, either individually or combinations thereof. In some variations, the cancer parameter may be predicted by an artificial intelligence algorithm including one or more of a statistical model, decision tree, rule-based system, simulation model, handcrafted algorithm, and machine learning model. In some variations, a cancer parameter may be predicted using one or more prediction machine learning models. For example, the cancer parameter may be predicted using a sensitivity-based prediction machine learning model and a second prediction using a specificity-based machine learning model. For example, the cancer parameter may be predicted using a first prediction machine learning model associated with a first stain or biomarker and a second prediction machine learning model associated with a second stain or biomarker. In some variations, the cancer parameter may include a cancer type such as one or more of non-cancer, colorectal cancer, lung cancer, pancreatic cancer, prostate cancer, breast cancer, esophageal cancer, liver cancer, ovarian cancer, kidney cancer, melanoma, or gastric cancer. In some variations, a human (e.g., pathologist, trained scientist) may analyze and score the morphological features of the ECMB data, as described in more detail with respect to FIGS.11A-11C and FIGS.50A-55 to predict a cancer parameter.Attorney Docket No.: AUMI-004 / 02WO 348385-2089
[0179] In some variations, the predicted cancer parameter may optionally be confirmed 112. For example, a first cancer parameter prediction may be confirmed by performing a second cancer parameter prediction using the same or different methodologies (e.g., stains, antibodies, prediction machine learning models, human analysis, fluid sample) in order to increase a confidence of the prediction. In particular, a first cancer parameter prediction may be analyzed using a first prediction machine learning model and a second cancer parameter prediction for confirming the first cancer parameter prediction may be analyzed using a second prediction machine learning model different from the first prediction machine learning model. Additionally or alternatively, a first cancer parameter prediction using a first a histology stain, for example, eosin stain or any combination of histology stains thereof may be confirmed with a second cancer parameter prediction using a second eosin stain or a colorectal cancer specific antibody biomarker to improve confidence in the cancer prediction. In some variations, a first cancer parameter prediction may be based on a technique not associated with ECMB analysis, such as a biopsy, and confirmed by a second cancer parameter prediction based on ECMB data (e.g., pathologist examining ECMB morphology, prediction machine learning model).
[0180] In some variations, predicting a cancer parameter may comprise a plurality of method described herein. Predicting a cancer parameter may include a first prediction by a human (e.g., a pathologist), a second prediction using radiomics-based image parameter analysis (e.g., extraction of quantitative image features using method described herein), and a third prediction using a prediction machine learning model. In some further variations, any combination, number, or sequence of these prediction methods may be used to confirm the cancer parameter or improve the confidence of the predicted cancer parameter.
[0181] As another example, a first cancer parameter prediction corresponding to a positive prediction for pre-cancer, a cancer type, and / or cancer stage (e.g., colorectal cancer, pre-cancer) based on pathologist analysis of morphology of the ECMB data may be considered in conjunction with a second cancer parameter prediction using a machine learning model. Alternatively, a first prediction based on a machine learning model may be considered in conjunction with a second prediction based on a pathologist analysis of the ECMB data. In this manner, the cancer prediction may be confirmed by analysis of the ECMB data using multiple techniques (e.g., human, machine learning, stains, antibodies, nucleic acids) to increase the confidence and / or confirm the prediction. Accordingly, the methods described herein mayAttorney Docket No.: AUMI-004 / 02WO 348385-2089 increase stakeholder (e.g., patient, pathologist, health care provider) confidence and utilization of the cancer parameter prediction. That is, a cancer prediction supported and / or confirmed by a prediction machine learning model, rather than solely reliant upon the prediction machine learning model, may increase stakeholder confidence in the cancer parameter prediction.
[0182] Additionally or alternatively, confirming the predicted cancer parameter may include analyzing ECMBs processed using different stains. For example, a first cancer parameter prediction may be based on ECMBs processed using a first stain (e.g., eosin stain) and a second (e.g., confirmatory) cancer parameter prediction may be based on ECMBs processed using a second stain (e.g., genetic stain highlighting DNA) different from the first stain. In some variations, a first cancer parameter prediction may be based on ECMBs processed using a first stain, and a second (e.g., confirmatory) cancer parameter prediction may be based on ECMBs processed using a combination of stains (e.g., mixture of IF antibody stains). Optionally, the combination of stains may include the first stain. In other variations, the a first cancer parameter prediction may be based on ECMBs processed using a stain or combination of stains, and second (e.g., confirmatory) cancer parameter prediction may be based on ECMBs processed without using stained (e.g., image using transmitted light microscopy) or vice versa. In some variations, a non-conclusive prediction may be refined through additional ECMB analysis. For example, a first cancer prediction may generically predict cancer based on first antibody stain. A second cancer parameter prediction may be based on ECMBs processed using a second antibody stain configured to identify a cancer type. In some variations, a first cancer parameter prediction may be configured to have relatively high sensitivity and low specificity while a second cancer parameter prediction may be configured to have a relatively low sensitivity and high specificity.
[0183] In some variations, confirming the predicted cancer parameter may include analyzing the processed ECMBs using one or more different prediction machine learning models each associated with one or more different stains. For example, a first cancer parameter prediction may be based on analysis by a first prediction machine learning model trained to analyze ECMBs processed using a first stain (e.g., eosin). A second (e.g., confirmatory) cancer prediction may be based on analysis by a second prediction machine learning model trained to analyze ECMBs process using a second stain (e.g., combination of IF stains). In some variations, a first prediction machine learning model associated with a first stain may be combined sequentially or in parallel with a second machine learning model associated with a differentAttorney Docket No.: AUMI-004 / 02WO 348385-2089 second stain. For example, the first and second model may be combined sequentially such that the second prediction machine learning model may analyze positive cancer predictions of the first prediction machine learning model to confirm the positive prediction of the first model. Such a sequential combination of prediction machine learning models may reduce the cost of performing cancer parameter prediction by reducing the number of ECMBs processed (e.g., processed with IF stains) for the second prediction machine learning model.
[0184] As another example, a first cancer parameter prediction based on pathologist analysis of morphology of the ECMB data as described herein in more detail may be confirmed by a second cancer parameter prediction using a prediction machine learning model. The prediction machine learning model may analyze the same morphology ECMB data as used in the pathologist analysis. For example, the prediction machine learning model may analyze a region of interest selected by the pathologist. Image parameters as described in more detail herein corresponding to the ECMB data of the region of interest may be generated and used by the prediction machine learning model to confirm the first cancer parameter. In some variations, a non-conclusive prediction by the pathologist may be refined through additional ECMB analysis by the prediction machine learning model. In some variations, one or more of the pathologist and the prediction machine learning model may segment the ECMB data for analysis. Additionally or alternatively, confirming the predicted cancer parameter using a prediction machine learning model may include analyzing ECMBs processed using different stains.
[0185] In some variations, one or more parameters of the second cancer parameter prediction (e.g., brightfield, stain, antibody biomarker, prediction machine learning model, human analysis) may be based on the first cancer parameter prediction. For example, a first cancer prediction may indicate features indicative of colorectal cancer and / or pancreatic cancer based on an eosin stain. A second cancer parameter prediction may be based on ECMBs processed using antibody biomarkers configured to differentiate between colorectal cancer and pancreatic cancer in order to confirm and / or refine the initial cancer parameter prediction.
[0186] In some variations, a first cancer parameter prediction (e.g., positive cancer prediction, pre-cancer) may initiate a second cancer parameter prediction based on a second ECMB sample of the subject different from a first ECMB sample used to generate the first cancer parameter prediction). For example, the first and second ECMB samples may come from different biological fluids (e.g., separate blood draws) of the subject or from different portions of theAttorney Docket No.: AUMI-004 / 02WO 348385-2089 biological fluids processed through different microfluidic devices. In some variations, the predicted cancer parameter may be confirmed a plurality of times, such as when the confirmation prediction does not match the initial prediction. In some variations, the predicted cancer parameter may be confirmed using a tissue sample (e.g., biopsy) or other methodology.
[0187] In some variations, analysis of the ECMB data may be confirmed using one or more of a pathologist and a prediction machine learning model. For example, the prediction machine learning model may be configured to identify one or more regions of interest used to predict a cancer parameter. A human (e.g., pathologist) may review these identified regions of interest and confirm that the prediction machine learning model regions of interest correspond regions of interest relevant to a cancer parameter, thereby increasing stakeholder confidence. Additionally or alternatively, regions of interest identified by a machine learning prediction model may be compared to regions of interest identified by a human (e.g., pathologist) to increase stakeholder confidence and prediction sensitivity. For example, when the regions of interest significantly differ, an analysis may be performed to determine if one or more of the human cancer parameter prediction and the prediction machine learning model cancer parameter prediction is in error.
[0188] In some variations, a confidence value of the predicted cancer parameter may be generated 114. For example, a confidence score may correspond to a probabilistic value assigned to a prediction, reflecting certainty or conviction of the prediction machine learning model prediction. For example, a confidence score may include a kappa score corresponding to reliability. Higher confidence scores correspond to greater trustworthiness of the predicted outcome, which is important in medical diagnosis. In some variations, confirmation of a cancer parameter prediction 112 may be based on one or more of a confidence score and a scoring rubric value. For example, a confidence score below a predetermined threshold (e.g., less than about 0.5, kappa score between about 50% and about 60%, kappa score between about 50% and about 55%, kappa score between about 55% and about 60%) may initiate a second cancer parameter prediction to confirm the first cancer parameter prediction. Similarly, a predetermined scoring rubric value (e.g., possible tumor 4, probable tumor 5, 6) performed by a human may initiate a second cancer parameter prediction. Alternatively, a control sample having a first cancer parameter prediction corresponding to a predetermined scoring rubric value (e.g., 2, 3) may initiate a second cancer parameter prediction.Attorney Docket No.: AUMI-004 / 02WO 348385-2089
[0189] In some variations, a prediction or suspected prediction of a cancer parameters may correspond to a lower confidence value (e.g., lower confidence score). A lower confidence score threshold may be desired in an initial screening to avoid classifying cancerous samples as false negatives. For example, a sensitivity-based prediction machine learning model may be used to generate a first cancer parameter prediction with a confidence score above a predetermined threshold (e.g., more than about 0.1, more than about 0.2, more than about 0.3, more than about 0.4, more than about 0.5, and more than about 0.55). The first cancer parameter prediction may be confirmed by second cancer parameter predication. A first cancer parameter prediction by a sensitivity-based prediction machine learning model may reduce the number of samples processed for a subsequent machine learning model (e.g., confirmatory prediction model, specificity-based model), thereby reducing the costs of predicting or confirming the cancer parameter. A high confidence score threshold may not correspond to high accuracy of a prediction machine learning model. In some variations, the confidence score threshold may be selected manually based on the performance of the trained model. For example, a machine learning model may predict a cancer (e.g., lung cancer) accurately at a default threshold (e.g., 0.5) and predict cancer with similar accuracy at a manually set at a lower threshold (e.g., 0.2). In some variations, a first prediction machine learning model with a lower threshold (e.g., sensitivity-based model) and a second prediction machine learning model with a high threshold (e.g., specificity-based model) may be used to predict a cancer parameter, thereby increasing the reliability of the prediction.
[0190] In some variations, a prediction or suspected prediction (e.g., positive for cancer) of a cancer parameter may correspond to a high confidence value. For example, a high confidence score threshold may be desired to avoid classifying non-cancerous samples as false positives. For example, a specificity-based prediction machine learning model may be used to generate a first cancer parameter prediction with a confidence score above a predetermined threshold (e.g., more than about 0.4, more than about 0.5, more than about 0.6, more than about 0.7, more than about 0.8, more than about 0.9, including all ranges and sub-values therebetween).
[0191] In some variations, the confidence score thresholds of two or more prediction machine learning models may be determined based on the performance of a combination of the two or more prediction machine learning models. For example, a first prediction machine learning model associated with a first stain (e.g., eosin-based model) may have a first confidenceAttorney Docket No.: AUMI-004 / 02WO 348385-2089 threshold (e.g., more than about 0.1, more than about 0.2, more than about 0.3, more than about 0.4, more than about 0.5, more than about 0.6, more than about 0.7, more than about 0.8, more than about 0.9). A second prediction machine learning model associated with a second stain (e.g., IF-based model) may have a second confidence threshold. The first model and the second model including their respective confidence score thresholds may be combined to predict cancer with high sensitivity, specificity, or overall accuracy based on a Youden’s J statistic. For example, a prediction with a confidence score above the predetermined first and second confidence score thresholds (e.g., more than about 0.2 and 0.7, more than about 0.3 and 0.4, more than about 0.3 and 0.7, more than about 0.4 and 0.7, more than about 0.5 and 0.6, more than about 0.6 and 0.5, more than about 0.7 and 0.2, more than about 0.7 and 0.3, more than about 0.7 and 0.5, including all ranges and sub-values therebetween) of the first and second machine learning model may correspond to a positive cancer prediction. In some variations, a confidence score threshold of one or more machine learning models or a confidence score of each machine learning model of a combination of two or more machine learning models may be optimized based on one or more of a Youden’s J statistic, specificity, and sensitivity associated with the model or combined model. In some variations, a confidence score threshold of one or more machine learning models or a confidence score of each machine learning model of a combination of two or more machine learning models or combination may be set manually.
[0192] In some variations, a confidence score threshold for one or more prediction machine learning models may be determined based on other methods or metric, instead of or in addition to, Youden’s J statistic. For example, the threshold may be selected based on: maximizing one or more of overall accuracy, sensitivity, specificity, precision, negative predictive value, and F1 score; using the point on the receiver operating characteristic (ROC) curve closest to the top-left corner of the plot (the “optimal operating point”); selecting the threshold that yields a desired tradeoff between sensitivity and specificity according to clinical requirements; setting the threshold according to a predetermined false positive rate or false negative rate; applying cost- sensitive or utility-based optimization; or utilizing cross-validation or bootstrapping to select the threshold that performs best on held-out data. In some variations, thresholds may be tuned using grid search, Bayesian optimization, or other algorithmic approaches. In other variations, thresholds may be set by expert review, regulatory guidelines, or consensus. Any combination of these thresholding methods, or other suitable thresholding techniques known in the art, may beAttorney Docket No.: AUMI-004 / 02WO 348385-2089 used to determine an appropriate confidence score threshold for a prediction machine learning model or models described herein.
[0193] In some variations, an electronic medical record (EMR) (e.g., electronic health record) may be updated 116. For example, the EMR may be updated to include one or more of the image data corresponding to biological fluid, the separated ECMBs, the ECMB data, the cancer parameter predictions, the generated confidence value, the methodologies used to generate the predictions (e.g., prediction machine learning model, pathologist scoring rubric), and the like. In some variations, a report may be generated for one or more stakeholders to include one or more of the image data corresponding to biological fluid, the separated ECMBs, the ECMB data, the cancer parameter predictions, the generated confidence value, the methodologies used to generate the predictions (e.g., prediction machine learning model, pathologist scoring rubric), the methodologies used to confirm the prediction, and the like. In some variations, one or more of the image data and ECMB data may be modified with one or more of a watermark (e.g., spatial domain, transform domain, region-based, reversible, authentication-based), copyright, steganography, and cryptography. For example, one or more of the image data and ECMB data may be packaged in a secure electronic format (e.g., CONTAINER, packaged electronic media) a digital watermark may facilitate secure electronic control and access (e.g., authentication, integrity verification) of subject data, which may facilitate licensing and copyright management as well as prevent manipulation, snooping, and other unauthorized access. Machine Learning Models
[0194] In some variations, a prediction machine learning model may be configured to output a cancer parameter prediction based on ECMB data generated by any of the systems, devices, and methods described herein. The prediction machine learning model may be trained using a training set of ECMB data based on transfer learning, ensemble learning, data augmentation, and deep learning techniques. FIG.3A depicts a flowchart representation of a method 300 of generating a prediction machine learning model. The method 300 may include receiving image data corresponding to ECMBs separated from a biological fluid 302 similar to step 102 of method 100 or step 9202 of method 9200, as described herein. For example, ECMB samples corresponding to one or more of stained non-cancerous ECMBs, breast cancer ECMBs, colorectal cancer ECMBs, prostate cancer ECMBs, lung cancer ECMBs, pancreatic cancer ECMBs, esophageal cancer ECMBs, liver cancer ECMBs, ovarian cancer ECMBs, kidneyAttorney Docket No.: AUMI-004 / 02WO 348385-2089 cancer ECMBs, melanoma ECMBs, or gastric cancer ECMBs may be received. In some variations, a prediction machine learning model may receive additional ECMB data such as biomarker data or non-ECMB data (e.g., an electronic medical record).
[0195] In some variations, the received image data may include images of one or more restriction channels of a microfluidic chip as described in more detail herein. For example, the prediction machine learning model may receive a first image of a first restriction channel (e.g., disposed along a perimeter) of the microfluidic chip and second image of a second restriction channel (e.g., disposed along a perimeter opposite the first restriction channel) of the microfluidic chip. In some variations, the prediction machine learning model may receive a combined image of the microfluidic chip including an image of both the first restriction channel and second restriction channel in the combined image. For example, the combined image may include the entire restriction region of the microfluidic chip. In some variations, the combined image may be configured such that it excludes an image of a middle channel (e.g., non- restriction channel) of the microfluidic chip. For example, a first image of a first restriction channel may be combined (e.g., appended, stitched) with a second image of a second restriction channel to produce a combined image of the restriction region with the middle channel and barriers (or any other microfluidic chip structure) removed. In some variations, the image data may comprise images including less than a full restriction channel (e.g., a portion of a restriction channel). For example, the microfluidic chip may be imaged by taking a plurality of images in a grid pattern to obtain image data from the chip. One or more images of the plurality may be combined (e.g., appended, stitched) to comprise the received image data. Combining a plurality of images may produce a larger and / or more detailed image to facilitate improved analysis of ECMBs separated on the chip. Additionally or alternatively, one or more of the images, including less than a full restriction channel (e.g., a portion of a restriction channel), may be processed separately as received image data. In some variations, the received image data may include one or more regions of interest and may be segmented into the one or more regions of interest as described herein in more detail.
[0196] In some variations, the received image data may be preprocessed 304. For example, the received image data may be cleaned, resized, normalized, and augmented. In some variations, data augmentation may include artificially expanding a training dataset of the received images by applying one or more transformations (e.g., rotations, flips, scaling, color adjustments) to theAttorney Docket No.: AUMI-004 / 02WO 348385-2089 original training dataset images, thereby generating new training images from the original training dataset, enriching the dataset diversity, and helping the prediction machine learning model learn from a broader range of sample variations. Preprocessing may reduce overfitting and provide robust model generalization. Data augmentation may include one or more of limited rotation, color adjustment, cropping, and noise injection to generate new training samples from the original training dataset to simulate a broader range of real-world variations. In some variations, an architecture of the prediction machine learning model may be designed 306. For example, the prediction machine learning model may comprise transfer learning including models such as one or more of ResNet (e.g., ResNet-50 pre-trained on ImageNet) and Inception machine learning models. Furthermore, the prediction machine learning models may incorporate regularization techniques such as dropout.
[0197] In some variations, the prediction machine learning model may be trained 308. For example, the prediction machine learning model may incorporate one or more of transfer learning (described in more detail with respect to method 400 and FIG.4), ensemble methods, and hyperparameter optimization. In some variations, the prediction machine learning model may be evaluated 310. For example, the performance of the prediction machine learning model in predicting one or more cancer parameters may be evaluated based on one or more of sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV) metrics, and area under the ROC curve (AUC). Additional or alternative methods of training the prediction machine learning model are described with reference to method 9200 of FIG.92.
[0198] FIG.4 depicts a flowchart representation of a method 400 of generating a prediction machine learning model incorporating transfer learning. The method 400 may include receiving image data corresponding to ECMBs separated from a biological fluid 402 similar to step 102 of method 100 and 302 of method 300. For example, ECMB samples corresponding to one or more of stained non-cancerous ECMBs, colorectal cancer ECMBs, prostate cancer ECMBs, lung cancer ECMBs, and pancreatic cancer ECMBs may be received. In some variations, the received image data may be preprocessed 404. For example, one or more morphological features (e.g., flakes, puncta, decorators) may be segmented in the received image data.
[0199] Furthermore, the received image data may be split for model training, validation, and testing. For example, FIG.5 is a plot 500 of a received image data set for a prediction machine learning model partitioned between training data, validation data, and testing data in order toAttorney Docket No.: AUMI-004 / 02WO 348385-2089 facilitate robust validation and independent testing. Moreover, the partitioning process may maintain a consistent distribution of cancer types across each of the training data, validation data, and testing data.
[0200] In some variations, an architecture of the prediction machine learning model may be designed 406 (e.g., ResNet-50). The prediction machine learning model may comprise a plurality of layers arranged to progressively refine the data representation using one or more techniques such as dropout (e.g., randomly ignoring units during training to prevent overdependence on any one feature) and batch normalization (e.g., standardizing inputs to each layer to stabilize and accelerate training) to improve feature extraction and classification. In some variations, the prediction machine learning model may comprise a plurality of convolutional layers configured to filter the input data. For example, the shallow convolutional layers may be configured to identify low-level features such as edges, textures, colors, and the like. As the data progresses through the neural network, the deeper convolutional layers may be configured to identify more complex patterns and structures corresponding to cancerous ECMBs. In some variations, one or more pooling layers may be interspersed between the convolutional layers in order to reduce the dimensionality of the feature maps. This down- sampling process may be configured to retain the most important features while reducing computational complexity, thereby enabling the prediction machine learning model to focus on the most relevant patterns. In some variations, one or more dropout layers may be coupled to one or more convolutional and fully connected layers, thereby teaching the model to learn redundant representations and enhancing its generalization capabilities. During training, a dropout layer may be configured to randomly ignore (e.g., sets to zero) a fraction of the neurons in these layers. For example, a dropout rate of 0.5 corresponds to randomly ignoring 50% of the neurons during each training iteration.
[0201] In some variations, hyperparameter optimization may comprise one or more grid search and randomized search methods. For example, grid search may be configured to explore a manually specified subset of the hyperparameter space of a prediction machine learning model. Randomized search may be configured to select random configurations and evaluate the performance of each. Hyperparameters may include one or more of learning rate (for neural networks), tree depth and number of estimators (for tree-based models), and dropout rate or batch size (for regularized deep learning models), tree depth, and the number of estimators inAttorney Docket No.: AUMI-004 / 02WO 348385-2089 bagging (e.g., number of models to train in parallel). The learning rate corresponds to the size of the steps taken during optimization to minimize the loss function.
[0202] In some variations, one more features may be extracted 408 (e.g., using deep learning techniques) to identify features that distinctly characterize cancerous ECMBs. Identifying one or more significant features for model training may optimize prediction machine learning model performance. For example, a final classification layer of a ResNet-50 model may be removed for deep feature extraction. In some variations, deep learning feature extraction may include using one or more of a convolutional neural network (CNNs) and feature selection techniques such as Principal Component Analysis (PCA), Recursive Feature Elimination (RFE), and feature importance regions generated by deep learning models. For example, CNNs may be configured to extract hierarchical features from the received image data where initial layers of the CNN identify low-level features (e.g., edges, textures) and deeper layers of the CNN identify high- level features (e.g., shapes, structures). PCA is a statistical technique that reduces the dimensionality of the data by transforming it into a set of orthogonal components that capture the directions of greatest variance. In this manner, PCA may support model training and identify features with the most variance within the data, thereby identifying features that may most significantly impact prediction machine model performance. RFE may be configured to iteratively remove less important features and retain the most influential features to improve cancer prediction. Feature importance regions may be identified using one or more of XRAI (eXtended Region Attention and Integration) and Integrated Gradients. Additional or alternative methods of feature selection are described with respect to image parameter selection in method 9200 of FIG.92 and may apply to feature selection for the prediction machine leaning model generally.
[0203] In some variations, a classifier may be added 410. For example, a classifier may comprise a global average pooling layer, a dense layer (e.g., 128-unit) with ReLU activation, a dropout layer configured to mitigate overfitting, and a softmax layer configured for binary classification.
[0204] In some variations, the prediction machine learning model may be trained 412. For example, the prediction machine learning model may be trained using one or more of a stochastic gradient descent (e.g., learning rate of 0.001), a cross-entropy loss function suitable for binary tasks, and one or more ensemble learning techniques including boosting 600 andAttorney Docket No.: AUMI-004 / 02WO 348385-2089 bagging 610. Furthermore, training may include an early stopping mechanism of over 30 epochs to optimize performance without overfitting. As shown in FIG.6, boosting 600 may include incremental training where multiple weak learners may be combined to generate a strong prediction machine learning model. For example, gradient boosting may be applied a plurality of times to refine model accuracy, particularly in classifying difficult instances. Gradient boosting refines models sequentially where each new model is trained to correct the residual errors of the combined ensemble of previous models, thereby effectively learning from errors and progressively improving the prediction capabilities of the model for challenging samples. In some variations, bagging 610 (e.g., bootstrap aggregating) may be configured to training multiple models in parallel on random subsets of the data (e.g., created with replacement), thereby ensuring diverse training scenarios (e.g., reducing variance without increasing bias) for each model within the ensemble. The final model predictions may be aggregated such as through a voting mechanism to produce a unified output.
[0205] In some variations, a prediction machine learning model may be trained on data corresponding to one or more of image data, ECMB data, clinical data, and any subset thereof. In some variations, the prediction machine learning model may be trained on synthetic data to improve a robustness of the prediction machine learning model and increase data available for testing and validation. For example, synthetic training data may be generated from manipulation of images of processed ECMBs by applying one or more transformations (e.g., rotations, flips, scaling, color adjustments) to the original training dataset images, thereby generating new training images from the original training dataset, enriching the dataset diversity, and helping the prediction machine learning model learn from a broader range of sample variations. In some variations, synthetic data may be created from existing ECMB image data sets by applying one or more of statistical methods (e.g., sampling, normalization, scaling, copula functions), generative machine learning methods (e.g., generative adversarial network, regression model), and data augmentation techniques (e.g., feature perturbation, noise injection), thereby generating new training samples from original training datasets to simulate a broader range of real-world variations. In some variations, a prediction machine learning model may be trained to analyze the ECMB data on unlabeled data. For example, a prediction machine learning model may be trained using unsupervised or semi-supervised techniques to learn patterns to ultimately distinguish between a cancerous and non-cancerous sample based on unlabeled data of one or more of a region of interest, a segment, a morphological feature, and an image parameter. SuchAttorney Docket No.: AUMI-004 / 02WO 348385-2089 techniques may be useful when limited labeled data is available to train a prediction machine learning model.
[0206] In some variations, a prediction machine learning model may be trained to segment received image data. Segments may correspond to portions of the received image data including ECMB data including one or more morphological features or image parameters. In some variations, a machine learning model may be trained to segment the received image data based on a manually identified region of interest.
[0207] In some variations, the prediction machine learning model may be evaluated based on one or more of sensitivity, specificity, PPV, and NPV metrics as shown in the performance metric equations 700 in FIG.7. Minimizing false negatives and providing consistent performance across a range of cancer types is important for clinical use. Morphology
[0208] In some variations, a cancer parameter may be predicted with high sensitivity and specificity using one or more of a prediction machine learning model and a human trained on ECMB morphology. For example, one or more morphological features of ECMB data may be segmented and one or more of a random forest machine learning model and a human may be configured to predict a cancer parameter such as pre-cancer, cancer type, and cancer stage. It may be helpful to briefly identify and describe the relevant ECMB morphology and microfluidic chip structure. As shown in image 210 of FIG.2B, the microfluidic chip may include a plurality of pillars 201 comprising a circular obstruction having, for example, a diameter of about 50 μm. The plurality of pillars 201 are within a fluid flow path but do not allow fluid to flow through the pillars 201. However, the pillars 201 may be configured such that biological material may adhere (e.g., immobilize) to a surface of the pillar 201. A background of the microfluidic chip may be defined as an area that allows fluid flow but does not include a pillar or biological material.
[0209] ECMB morphology may be classified generally as one or more of flakes, puncta, and decorators. FIGS.8A-8D are images 800-895 of ECMBs including flakes. As shown in FIG.8A, the flakes may comprise eosinophilic (pink) structures 802, 812, 822, 832, 842, 852, 854, 862, 872, 882, 884, 892, 894, 896 having an area of between about 5 μm2and about 450 μm2, about 5 μm2and about 400 μm2, about 5 μm2and about 300 μm2, about 5 μm2and about 200 μm2, aboutAttorney Docket No.: AUMI-004 / 02WO 348385-2089 100 μm2and about 450 μm2, about 200 μm2and about 450 μm2, about 300 μm2and about 450 μm2, about 200 μm2and about 300 μm2, about 5 μm2and about 1000 μm2including all ranges and sub-values in-between.
[0210] In some variations, flakes may be classified morphologically as one or more of crystals, bands, and veils. For examples, crystals 854 may comprise paracrystalline sharply angulated structures, bands 884 may comprise flakes with smoothly rounded contours, and veils 862 may comprise flakes with curvilinear folds. Furthermore, flakes may be classified based on eosin staining characteristics including translucency (e.g., translucent 872, non-translucent 882) and color (e.g., deep eosinophilic, pale eosinophilic, gray). Flakes may be classified architecturally as one or more of molding, touching, free, and bridging. For example, molding 894 may comprise flakes partially or completely encircling a pillar, touching 852 may comprise flakes in contact with a pillar without molding or partial encirclement, free 896 may comprise flakes that do not touch a pillar, and bridging 892 may comprise flakes that mold to adjacent pillars. The flakes shown in FIG.8B correspond to human plasma stained with 0.125% w / v Eosin in a 15 μm depth microfluidic channel. Pillars p have a diameter of about 50 μm and are located in a microfluidic channel of a microfluidic chip as described herein.
[0211] FIGS.9A and 9B are images 900-950 of ECMBs including puncta where the puncta is not in contact with a pillar and the punctum comprises a circular structure 902, 904, 912, 932, 942, 944, 946, 948, 952, 954, 956, 958, 960 having an area less than about 15 μm2, less than about 10 μm2, less than about 5 μm2, between about 1 μm2and about 15 μm2, between about 1 μm2and about 15 μm2, between about 3 μm2and about 14 μm2, between about 5 μm2and about 10 μm2, between about 1 μm2and about 10 μm2, and between about 5 μm2and about 15 μm2, including all ranges and sub-values in-between.
[0212] In some variations, punctum may be classified morphologically as one or more granules and droplets. For example, granules may comprise structures with angulated paracrystalline morphology (e.g., sharply demarcated border), and droplets may comprise an ovoid (e.g., rounded) structure that may include smudgy borders. Furthermore, punctum may be classified based on eosin staining characteristics including translucency (e.g., translucent 954, 956, non-translucent 952, 958, 960) and color (e.g., deep eosinophilic 944, pale eosinophilic 946, gray 942). Puncta may be classified architecturally as one or more of a string 948 and a speckle. For example, a string may comprise a pearl on a string-like arrangement of puncta.Attorney Docket No.: AUMI-004 / 02WO 348385-2089 Some strings may be disposed in a background of faintly eosinophilic strand-like material. Spreckles may comprise a speckled, haphazard architectural arrangement.
[0213] FIGS.10A-10C are images 1000-1070 of ECMBs including decorators having one or more circular structures 1002, 1012, 1022, 1032, 1042, 1044, 1046, 1052, 1054, 1062, 1072 (e.g., granule) adjacent to a pillar 1001. The decorators may have an area of less than about 7 μm2, between about 2 μm2and about 40 μm2, between about 2 μm2and about 35 μm2, between about 2 μm2and about 20 μm2, between about 2 μm2and about 10 μm2, and between about 2 μm2and about 7 μm2, including all ranges and sub-values in-between.
[0214] In some variations, decorators may be classified morphologically as one or more granules and droplets. For example, granules 1072 may comprise structures with angulated paracrystalline morphology (e.g., sharply demarcated border), and droplets 1062 may comprise an ovoid (e.g., rounded) structure that may include smudgy borders. Furthermore, decorators may be classified based on eosin staining characteristics including translucency (e.g., translucent 1054, non-translucent 1052) and color (e.g., deep eosinophilic 1044, pale eosinophilic 1046, gray 1042). Decorators may be classified architecturally based on the number degrees in contact with the pillars. Morphological Features
[0215] In some variations, a cancer parameter may be predicted based on a morphological feature using one or more of a prediction machine learning model and human analysis of ECMB morphology. Generally, separated ECMBs (e.g., matricles) may include morphological features that may be used to distinguish cancerous samples from non-cancerous samples. As described in more detail herein, a morphological feature may correspond to one or more of the geometry (e.g., shape, size) of the ECMB, spatial location of the ECMB, intensity of the ECMB and / or biomarkers, abundance of the ECMB and / or biomarkers, color of the ECMB, texture characteristics (e.g., contrast, entropy, or uniformity) of the ECMB, edge properties (e.g., sharpness or boundary definition) of the ECMB, stiffness, and other biomechanical attributes of the ECMB. For example, a morphological feature corresponding to the geometry the ECMBs may include one or more dimensions such length, width, height, area, angles, and curvatures. A morphological feature corresponding to color may be based on the consistency and variation in the stain color of the ECMB or intensity of stain at one or more localizations in the ECMB. AAttorney Docket No.: AUMI-004 / 02WO 348385-2089 morphological feature corresponding to an edge property of an ECMB may include sharp edges or transitions between the ECMB and background or contrast. Other morphological features may correspond to the directionality, spatial relationships, frequency, orientation, or brightness of the ECMB. The background may include a microfluidic chip and one or more pillars. In some variations, the morphological features may form in the ECMBs as a result of processing on a microfluidic chip. In some variations, a morphological feature may be more strongly associated with a first cancer type compared to second cancer type.
[0216] A fluid sample associated with a cancer may include mostly healthy looking ECMBs but may be predicted as cancerous based on a relatively few ECMBs corresponding to cancer. In some variations, a positive cancer prediction or cancer stage prediction may be based on less than all the ECMBs in a biological fluid sample. For example, a cancer parameter may be predicted based on less than about 20% of ECMBs in a biological fluid sample, less than about 10% of ECMBs in a biological fluid sample, less than about 5% of ECMBs in a biological fluid sample, and less than about 1% of ECMBs in a biological fluid sample. A processed biological fluid sample may produce ECMBs including morphological features associated with cancer and non-cancer. For example, a first ECMB of a fluid sample indicative of cancer and a second ECMB of the same fluid sample indicative of non-cancer may be differentiated by a morphological feature expressed differently in the first ECMB compared to the second ECMB. The entire fluid sample may be predicted as cancerous based on the first ECMB. In some variations, a region of interest corresponding to a portion of a larger image of a biological fluid sample may be selected to include one or more ECMBs. The region of interest may be selected to include ECMBs associated with cancer or suspected of being associated with cancer as determined by one or more of a pathologist and predictive machine learning model. A cancer parameter prediction may be based on one or more morphological features in a region of interest.
[0217] In some variations, a morphological feature including an edge property of an ECMB and may be used to predict a cancer parameter. For example, FIG.37 includes images 3710 and 3720 of ECMBs differentiated by a morphological feature corresponding to an edge property of the ECMB. In some variations, an ECMB indicative of cancer may express a relatively less defined boundary and lower change in intensity of the image between the foreground (e.g., stained ECMB) and background (e.g., microfluidic chip and pillars) compared to a non- cancerous ECMB. For example, image 3710 shows a pillar 3714 and an ECMB 3712 associatedAttorney Docket No.: AUMI-004 / 02WO 348385-2089 with pancreatic cancer. The cancer associated ECMB 3712 may be characterized by an edge that is rough, crumbly, and poorly defined (e.g., not sharp). By contrast, image 3720 shows a pillar 3724 and an ECMB 3722 associated with a non-cancerous sample. The non-cancer associated ECMB 3722 may be characterized by an edge that is smooth, sharply contrasted, and well- defined with clear borders.
[0218] In some variations, a morphological feature including a stiffness of an ECMB may be used to predict a cancer parameter. FIG.38 includes images 3810 and 3820 of ECMBs differentiated by a morphological feature corresponding to a stiffness of the ECMB. In some variations, an ECMB associated with cancer may include relatively fewer folds compared to a non-cancerous ECMB which may include more folds when contacting a pillar. For example, image 3810 shows a pillar 3814 and an ECMB 3812 associated with pancreatic cancer. The cancer associated ECMB 3812 may be characterized by a stiffness and general rigidity, minimal folding, and a resistance to changing form when contacting a pillar 3814. By contrast, image 3820 shows a pillar 3824 and an ECMB 3822 associated with a non-cancerous sample. The non- cancer associated ECMB 3822 may be characterized by a collapse in the ECMB at the pillar 3824 and a relatively higher number of folds.
[0219] In some variations, a morphological feature including a heterogeneity in staining color of an ECMB may be used to predict a cancer parameter. FIG.39 includes images 3910 and 3920 of ECMBs differentiated by a morphological feature corresponding to a heterogeneity in staining color of the ECMB. The heterogeneity of staining color of the ECMB may be based on a presence or an absence of a color change in the ECMB. For example, the color change of the ECMB may include a change in one or more a hue, a saturation, and a brightness (e.g., value). Hue may generally refer to the dominant wavelength of light given off by the sample (e.g., red, blue, green). Saturation may generally refer to the purity of a color. For example, an example of high saturation may include strong, vivid coloring and an example of low saturation may include muted coloring. Brightness may generally refer to how light or dark the coloring appears. An ECMB associated with cancer may include relatively greater color changes (e.g., patterns, high heterogeneity) compared to a healthy ECMB which may include relatively less color change (e.g., uniform coloring). For example, image 3910 shows a pillar 3914 and an ECMB 3912 associated with pancreatic cancer. The cancer associated ECMB 3912 may be characterized by heterogeneity in coloring. By contrast, image 3920 shows a pillar 3924 an ECMB 3922Attorney Docket No.: AUMI-004 / 02WO 348385-2089 associated with a non-cancerous sample. The non-cancer associated ECMB 3922 may be characterized by homogeneity in coloring.
[0220] In some variations, a morphological feature including a texture of an ECMB may be used to predict a cancer parameter. FIG.40 includes images 4010 and 4020 of ECMBs differentiated by a morphological feature corresponding to a texture of the ECMB. In some variations, an ECMB associated with cancer may be granular and have relatively complex arrangements (e.g., structures) of internal material compared to a healthy ECMB having relatively minimal granularity and a smoother appearance. For example, image 4010 shows a pillar 4014 and an ECMB 4012 associated with pancreatic cancer. The cancer associated ECMB 4012 may be characterized as highly variable in internal structure, granular in appearance, bubble-like, and highly textured. By contrast, image 4020 shows a pillar 4024 an ECMB 4022 associated with non-cancer. The non-cancer associated ECMB 4022 may be characterized as smooth, flat, punctuated, and less textured.
[0221] In some variations, a morphological feature including a color of an ECMB may be used to predict a cancer parameter. FIG.41 includes images 4110 and 4120 of ECMBs differentiated by a morphological feature corresponding to a flesh-color staining of the ECMB. For example, an ECMB associated with a cancer stained with a predetermined concentration of eosin may have a flesh tone stain color (e.g., salmon color) compared to a healthy ECMB stained with the same predetermined concentration of eosin having a dark pink (e.g., eosinophilic color) coloring. For example, image 4110 shows a pillar 4114 and an ECMB 4112 associated with pancreatic cancer. The cancer associated ECMB 4112 may be characterized as tan and / or flesh color stained, and having staining that is not fully eosinophilic. By contrast, image 4120 shows a pillar 4124 an ECMB 4122 associated with a non-cancerous sample. The non-cancer associated ECMB 4122 may be characterized as eosinophilic and pink in color.
[0222] In some variations, a cancer parameter including a cancer stage may be predicted based on one or more morphological features. Generally, one or more morphological features present in the ECMBs of a biological fluid sample may associate the sample with a stage of cancer (e.g., stage I pancreatic cancer, stage IV pancreatic cancer). In some variations, a cancer stage prediction may be based on a degree of a difference between a first sample and a second control sample (e.g., known non-cancer sample) in one or more morphological features. For example, a relatively slight difference in a morphological feature corresponding to coloring between a firstAttorney Docket No.: AUMI-004 / 02WO 348385-2089 ECMB sample and a non-cancerous ECMB sample may correspond to a prediction of an earlier stage (e.g., stage I) cancer and a significant difference in the morphological feature (e.g., coloring) may correspond to a prediction of a later stage (e.g., stage IV) of cancer. Additionally or alternatively, a cancer stage prediction may be based on a number of morphological features associated with cancer included in the ECMBs of a biological fluid sample. For example, the presence of many morphological features associated with cancer may correspond to a prediction of a later stage of cancer while the presence of fewer morphological features associated with cancer may correspond to a prediction of an earlier stage of cancer. As part of cancer management and treatment, the stage of cancer can be a critical factor in determining prognosis and treatment strategies. In this context, a morphological feature or features may be used to detect minimal residual disease (MRD), which may indicate a patient’s risk for relapse after a course of treatment, or to monitor the effectiveness of a therapeutic regimen. As explained herein, a positive cancer prediction or cancer stage prediction may be based on a less than all the ECMBs in a biological fluid sample.
[0223] In some variations, a cancer parameter including a cancer stage (including pre-cancer) may be predicted using a prediction machine learning method. In some variations, the prediction machine learning method may predict a cancer stage based on one or more morphological features and / or image parameters as described in more detail herein. For example, the ability to detect and stage pancreatic ductal adenocarcinoma (PDAC), as described with respect to FIGS. 42A-45B, may provide critical diagnostic and prognostic information for clinical management of the disease. Accurate characterization of early-stage disease (e.g., stage I, II) is crucial as it enables timely intervention and improves patient outcomes. The accurate staging of advanced disease (e.g., stage III and IV) is also vital for informing prognosis.
[0224] For example, FIGS.42A and 42B are images 4200a-4230a and 4200b-4230b of ECMBs associated with stage I pancreatic cancer (e.g., PDAC I) and a non-cancerous sample (e.g., healthy control). Image 4200a shows a processed biological fluid of a 66-year-old female and identifies three regions of interest 4210a, 4220a, 4230a including ECMBs associated with a stage I pancreatic cancer. Image 4200b shows a processed biological fluid of another 66-year-old female and identifies three regions of interest 4210b, 4220b, 4230b including ECMBs associated with a non-cancerous sample. Both biological fluid samples include about 200 µl of plasma stained with eosin (0.125% w / v) processed at about 100 mmHg. Regions of interest 4210a,Attorney Docket No.: AUMI-004 / 02WO 348385-2089 4220a, and 4230a include one or more pillars 4242 and ECMBs 4212a, 4222a, and 4232a having morphological features associated with cancer. By contrast, regions of interest 4210b, 4220b, and 4230b include one or more pillars 4242 and ECMBs 4212b, 4222b, and 4232b having morphological features associated with a non-cancerous sample.
[0225] ECMB 4212a may be differentiated from ECMB 4212b based on morphological features corresponding to an edge of the ECMB and a stiffness of the ECMB. ECMB 4222a may be differentiated from ECMB 4222b based on morphological features corresponding to an edge of the ECMB, staining pattern of the ECMB, and texture of the ECMB. ECMB 4232a may be differentiated from ECMB 4232b based on a morphological feature corresponding to a flow pattern of the ECMB. For example, ECMB 4232a may be characterized by a nonlinear flow pattern, while ECMB 4232b may be characterized by a linear flow pattern. The flow pattern may correspond to a difference in viscosity between cancerous and non-cancerous samples. Thus, a cancer parameter of stage I cancer may be predicted for the sample of FIG.42A.
[0226] FIGS.43A and 43B are images 4300a-4330a and 4300b-4330b of ECMBs associated with stage II pancreatic cancer (e.g., PDAC II) and a non-cancerous sample (e.g., healthy control). Image 4300a shows a processed biological fluid of a 65-year-old female and identifies three regions of interest 4310a, 4320a, 4330a including ECMBs associated with stage II pancreatic cancer. Image 4300b shows a processed biological fluid of another 65-year-old female and identifies three regions of interest 4310b, 4320b, 4330b including ECMBs associated with a non-cancerous sample. Both biological fluid samples include about 200 µl of plasma stained with eosin (0.125% w / v) processed at about 100 mmHg. Regions of interest 4310a, 4320a, and 4330a include one or more pillars 4342 and ECMBs 4312a, 4322a, and 4332a having morphological features associated with cancer. By contrast, regions of interest 4310b, 4320b, and 4330b include one or more pillars 4342 and ECMBs 4312b, 4322b, and 4332b having morphological features associated with a non-cancerous sample. In some variations, ECMB 4312a may be differentiated from ECMB 4312b based on morphological features corresponding to an edge of the ECMB, stiffness of the ECMB, and texture of the ECMB. Similarly, ECMB 4322a may be differentiated from ECMB 4322b based on morphological features corresponding to an edge of the ECMB, staining pattern of the ECMB, stiffness of the ECMB, and texture of the ECMB. ECMB 4332a may be differentiated from ECMB 4332b based on morphological features corresponding to an edge of the ECMB, staining pattern of the ECMB, stiffness of theAttorney Docket No.: AUMI-004 / 02WO 348385-2089 ECMB, and texture of the ECMB. The amount of differences in morphological features between the samples of FIGS.43A and 43B may be greater than the amount of differences in morphological features between the samples of FIGS.42A and 42B such that a cancer parameter of stage II cancer may be predicted for the sample of FIG.43A.
[0227] FIGS.44A and 44B are images 4400a-4430a and 4400b-4430b of ECMBs associated with stage III pancreatic cancer (e.g., PDAC III) and a non-cancerous sample (e.g., healthy control). Image 4400a shows a processed biological fluid of a 50-year-old male and identifies three regions of interest 4410a, 4420a, 4430a including ECMBs associated with stage III pancreatic cancer. Image 4400b shows a processed biological fluid of another 50-year-old male and identifies three regions of interest 4410b, 4420b, 4430b including ECMBs associated with a non-cancerous sample. Both biological fluid samples include about 200 µl of plasma stained with eosin (0.125% w / v) processed at about 100 mmHg. Regions of interest 4410a, 4420a, and 4430a include one or more pillars 4442 and ECMBs 4412a, 4422a, and 4432a having morphological features associated with cancer. By contrast, regions of interest 4410b, 4420b, and 4430b include one or more pillars 4442 and ECMBs 4412b, 4422b, and 4432b having morphological features associated with a non-cancerous sample. In some variations, ECMB 4412a may be differentiated from ECMB 4412b based on morphological features corresponding to an edge of the ECMB, staining pattern (e.g., heterogeneity) of the ECMB, and texture of the ECMB. Next, ECMB 4422a may be differentiated from ECMB 4422b based on morphological features corresponding to an edge of the ECMB, staining pattern of the ECMB, stiffness of the ECMB, and texture of the ECMB. Similarly, ECMB 4432a may be differentiated from ECMB 4432b based on morphological features corresponding to an edge of the ECMB, staining pattern of the ECMB, and texture of the ECMB. The difference in one or more morphological features between the samples of FIGS.44A and 44B may be characterized as more than the difference in morphological features between the samples of FIGS.43A and 43B such that a cancer parameter of stage III cancer may be predicted for the sample of FIG.44A.
[0228] FIGS.45A and 45B are images 4500a-4530a and 4500b-4530b of ECMBs associated with stage IV pancreatic cancer (e.g., PDAC IV) and a non-cancerous sample (e.g., control). Image 4500a shows a processed biological fluid of a 64-year-old male and identifies three regions of interest 4510a, 4520a, 4530a including ECMBs associated with stage IV pancreatic cancer. Image 4500b shows a processed biological fluid of another 64-year-old male andAttorney Docket No.: AUMI-004 / 02WO 348385-2089 identifies three regions of interest 4510b, 4520b, 4530b including ECMBs associated with a non- cancerous sample. Both biological fluid samples include about 200 µl of plasma stained with Eosin (0.125% w / v) processed at about 100 mmHg. Regions of interest 4510a, 4520a, and 4530a include one or more pillars 4542 and ECMBs 4512a, 4522a, and 4532a having morphological features associated with cancer. By contrast, regions of interest 4510b, 4520b, and 4530b include one or more pillars 4542 and ECMBs 4512b, 4522b, and 4532b having morphological features associated with a non-cancerous sample. ECMB 4512a may be differentiated from ECMB 4512b based on morphological features corresponding to an edge of the ECMB, staining pattern (e.g., heterogeneity) of the ECMB, stiffness of the ECMB, and texture of the ECMB. Next, ECMB 4522a may be differentiated from ECMB 4522b based on morphological features corresponding to an edge of the ECMB, staining pattern of the ECMB, stiffness of the ECMB, texture of the ECMB, and flesh coloring of the ECMB. Finally, ECMB 4532a may be differentiated from ECMB 4532b based on morphological features corresponding to an edge of the ECMB, staining pattern of the ECMB, and texture of the ECMB. The difference in one or more morphological features between the samples of FIGS.45A and 45B may be characterized as more numerous than the difference in morphological features between the samples of FIGS.44A and 44B such that a cancer parameter of stage IV cancer may be predicted for the sample of FIG.45A.
[0229] The ability to differentiate between the morphology of cancerous, pre-cancerous, and non-cancerous ECMBs may enable the development of robust methods capable of accurately classifying cancer from control samples across all disease stages and various cancer types. Furthermore, this morphological approach may provide the detail necessary to differentiate between disease stages and cancer types, thereby offering a powerful tool not only in diagnostic but also in prognostic evaluation and monitoring.
[0230] For illustrative purposes, morphological features associated with cancer have been compared to similar morphological features associated with a non-cancerous sample. However, a cancer parameter including one or more of pre-cancer, a cancer type, and a cancer stage may be predicted without comparison to a non-cancerous control sample. For example, an ECMB morphology trained pathologist and / or prediction machine learning model may analyze ECMB data without reference to a non-cancerous morphology (e.g., healthy control reference). In some variations a cancer parameter including a first cancer type may be predicted with comparison to samples of a second cancer type.Attorney Docket No.: AUMI-004 / 02WO 348385-2089
[0231] As an example of morphological features useful in the prediction, diagnosis, and / or prognosis of additional cancer types, FIGS.107 and 108 depict ECMB samples of subjects with varying stages of gastric cancer, separated on a microfluidic chip. To produce the images 10710, 10720, 10810, 10820, 10830, 10840, plasma samples were acquired prior to chemotherapeutic intervention from a cohort of 45-80 year-old patients. Samples from patients with a history of malignancy were excluded. Plasmas were collected in K2 EDTA tubes and subject to centrifugation from 1000g-3000g. To isolate extracellular matrix bodies (ECMBs), plasma was perfused through the chip, followed by washing with PBS. Images were acquired using brightfield microscopy at 10x using Axioscan 7 and an Axiocam 705 color camera.
[0232] FIG.107 are images 10710, 10720 of ECMBs associated with gastric cancer. Image 10710 shows a processed biological fluid including ECMBs associated with stage II gastric cancer. Image 10720 shows a processed biological fluid including ECMBs associated with stage IV gastric cancer. Images 10710 and 10720 include one or more pillars and indicated ECMBs (see arrows) that may be differentiated from healthy control ECMBs or ECMBs associated with other cancer types based on morphological features corresponding to a ribbon structure of the ECMBs. The indicated gastric cancer ECMBs of images 10710, 10720 may be distinct from control ECMBs in that ribbon structures were common. Ribbon structures may be defined as helical fibers about 10 µm to about 100 µm long or about 10 µm to about 50 µm long, including all subranges and values therein, such as about 30 µm long. FIG.108 are images 10810, 10820, 10830, 10840 of ECMBs associated with gastric cancer. Image 10810 shows a processed biological fluid including ECMBs associated with stage I gastric cancer. Image 10820 shows a processed biological fluid including ECMBs associated with stage III gastric cancer. Image 10830 shows another processed biological fluid including ECMBs associated with stage III gastric cancer. Image 10840 shows another processed biological fluid including ECMBs associated with stage I gastric cancer. Images 10810, 10820, 10830, 10840 include one or more pillars and indicated ECMBs that may be differentiated from healthy control ECMBs or ECMBs associated with other cancer types based on morphological features corresponding to nodules of the ECMBs. Nodules may be circular or generally ovoid bodies with strong borders about 10 µm to about 50 µm wide, including all subranges and values therein. In gastric cancer ECMBs, nodules may be organized within larger structures more often compared to healthy control ECMBs. In some variations, rigid larger ECMBs were also common in gastric cancer plasma compared to healthy control plasma. The rigidness (e.g., stiffness) of the ECMB may beAttorney Docket No.: AUMI-004 / 02WO 348385-2089 indicated by a lack of wrapping interactions with the pillars.As another example of morphological features useful in the prediction, diagnosis, and / or prognosis of additional cancer types, FIGS.109 and 110 depict ECMB samples of subjects with varying stages of esophageal cancer, separated on a microfluidic chip. To produce the images 10910, 10920, 11010, 11020, 11030, 11040, plasma samples were acquired prior to chemotherapeutic intervention from a cohort of 45-80 year-old patients. Samples from patients with a history of malignancy were excluded. Plasmas were collected in K2 EDTA tubes and subject to centrifugation from 1000- 3000g. To isolate extracellular matrix bodies (ECMBs), plasma was perfused through the chip, followed by washing with PBS. Images were acquired using brightfield microscopy at 10x using Axioscan 7 and an Axiocam 705 color camera.
[0233] FIG.109 are images 10910, 10920 of ECMBs associated with esophageal cancer. Image 10910 shows a processed biological fluid including ECMBs associated with stage III esophageal cancer. Image 10920 shows a processed biological fluid including ECMBs associated with stage II esophageal cancer. Images 10910 and 10920 include one or more pillars and indicated ECMBs that may be differentiated from healthy control ECMBs based on morphological features corresponding to a ribbon structure of the ECMBs. The indicated esophageal cancer ECMBs of images 10910, 10920 may be distinct from control ECMBs in that ribbon structures were common. Ribbon structures may be defined as helical fibers about 10 µm to about 100 µm long or about 10 µm to about 50 µm long, including all subranges and values therein, such as about 30 µm long. FIG.110 are images 11010, 11020, 11030, 11040 of ECMBs associated esophageal cancer. Image 11010 shows a processed biological fluid including ECMBs associated with stage III esophageal cancer. Image 11020 shows a processed biological fluid including ECMBs associated with stage II esophageal cancer. Image 11030 shows another processed biological fluid including ECMBs associated with stage II esophageal cancer. Image 11040 shows another processed biological fluid including ECMBs associated with stage III esophageal cancer. Images 11010, 11020, 11030, 11040 include one or more pillars and indicated esophageal ECMBs that may be differentiated from healthy control ECMBs or ECMBs associated with other cancer types based on morphological features corresponding to nodules of the ECMBs. Nodules may be circular or generally ovoid bodies with strong borders about 10 µm to about 50 µm wide, including all subranges and values therein. In some variations, rigid larger ECMBs were also common in esophageal cancer plasma compared toAttorney Docket No.: AUMI-004 / 02WO 348385-2089 healthy control plasma. The rigidness (e.g., stiffness) of the ECMB may be indicated by a lack of wrapping interactions with the pillars.
[0234] As yet another example of morphological features useful in the prediction, diagnosis, and / or prognosis of additional cancer types, FIGS.111, 112, and 113 depict ECMB samples of subjects with varying stages of liver cancer, separated on a microfluidic chip. To produce the images 11110, 11120, 11130, 11140, 11150, 11210, 11220, 11230, 11310, 11320, 11330, 11340, plasma samples were acquired prior to chemotherapeutic intervention from a cohort of 45-80 year-old patients. Samples from patients with a history of malignancy were excluded. Plasmas were collected in K2 EDTA tubes and subject to centrifugation from 1000-3000g. To isolate extracellular matrix bodies (ECMBs), plasma was perfused through the chip, followed by washing with PBS. Images were acquired using brightfield microscopy at 10x using Axioscan 7 and an Axiocam 705 color camera.
[0235] FIG.111 are images 11110, 11120, 11130, 11140, 11150 of ECMBs associated with liver cancer. Image 11140 shows a processed biological fluid including ECMBs associated with stage II liver cancer. Images 11110, 11150 show processed biological fluids including ECMBs associated with stage III liver cancer. Image 11120, 11130 show processed biological fluids including ECMBs associated with stage IV liver cancer. Images 11110, 11120, 11130, 11140, 11150 include one or more pillars and indicated liver cancer ECMBs (see arrows) that may be differentiated from healthy control ECMBs or ECMBs associated with other cancer types based on morphological features corresponding grainy ECMBs. The indicated esophageal cancer ECMBs of images 11110, 11120, 11130, 11140, 11150 may be distinct from control ECMBs in that grainy ECMBs were common. Grainy ECMBs may be defined as large pillar-spanning structures comprised of dense uniform granular material about 1 µm to about 5 µm or about 1 µm to about 3 µm, including all subranges and values therein and having a “granite-like” structure.
[0236] FIG.112 are images 11210, 11220, 11230 of ECMBs associated with liver cancer. Image 11230 shows a processed biological fluid including ECMBs associated with stage I liver cancer. Image 11220 shows a processed biological fluid including ECMBs associated with stage II liver cancer. Image 11210 shows a processed biological fluid including ECMBs associated with stage III liver cancer. Images 11210, 11220, 11230 include one or more pillars and indicated liver cancer ECMBs that may be differentiated from healthy control ECMBs orAttorney Docket No.: AUMI-004 / 02WO 348385-2089 ECMBs associated with other cancer types based on morphological features corresponding to small plaques of the ECMBs. As indicated in images 11210, 11220, 11230 small plaques may be lighter in color compared to surrounding ECMBs or healthy control ECMBs.
[0237] FIG.113 are images 11310, 11320, 11330, 11340 of ECMBs associated with liver cancer. Images 11310, 11320, 11330 show processed biological fluids including ECMBs associated with stage II liver cancer. Image 11340 shows a processed biological fluid including ECMBs associated with stage IV liver cancer. Images 11310, 11320, 11330, 11340 include one or more pillars and processed liver cancer ECMBs that may be differentiated from healthy control ECMBs or ECMBs associated with other cancer types based on morphological features corresponding to pleomorphic ECMBs and / or heterogeneity of the ECMB. For example, separated liver cancer ECMBs may display a greater variety of structures and shapes compared to healthy control ECMBs. Image Parameters
[0238] In some variations, image parameters may be generated from ECMB image data corresponding to separated ECMBs and used to predict one or more cancer parameters. The image parameters may correspond to quantifiable features (e.g., numeric values) of an image of processed ECMBs or region of interest selected from the image. In this manner, the image parameters, in some variations, may correspond with and / or quantify a morphological feature. Image parameters may be considered ECMB data and may be analyzed by a prediction machine learning model to predict a cancer parameter. Prediction machine learning models and / or pathologists may predict a cancer parameter with high sensitivity and specificity based on image parameters and, optionally, other ECMB data. For example, a prediction machine learning model may be trained to classify pre-cancer, early-stage cancer, and / or late-stage cancer from non- cancerous samples for a variety of cancer types (e.g., PDAC, CRC, AA). In some variations, image parameters may improve transparency of a prediction machine learning model by providing an explanation of what information in an image the model is emphasizing when predicting a cancer parameter.
[0239] Generally, the image parameters may be associated with one or more of intensity features of an image (e.g., color channel intensity ranging from values of 0-255), color features of the image, textures of the image (e.g., Gabor image parameters, gray level cooccurrenceAttorney Docket No.: AUMI-004 / 02WO 348385-2089 matrix (GLCM)), edge features of the image, wavelet features of the image, statistical parameters, and / or any radiomic features known in other medical imaging applications (e.g., MRI, CT, or PET scans). For example, a plurality of image parameters may include one or more of a green image channel mean, a mean intensity (e.g., quantification of overall brightness), intensity skewness, contrast, homogeneity, low frequency energy, high frequency energy, correlation, edge entropy, standard deviation, Gabor features, wavelet features, and / or kurtosis. In some variations, a prediction machine learning model may discriminate between cancerous and non-cancerous samples based on morphological distinction in the ECMBs captured by one or more image parameters.
[0240] A cancer parameter may be predicted based on ECMB data including one or more image parameters. In some variations, a cancer parameter prediction may be confirmed using a prediction machine learning model analyzing image parameters. In some variations, a prediction machine learning model may analyze image parameters to confirm the association of a morphological feature with cancer or non-cancer. For example, FIG.46A is an illustrative plot 4610 of mean intensity curves for 25 pancreatic cancer samples, and FIG.46B is an illustrative plot 4620 of mean intensity curves for 25 non-cancerous samples. Plots 4610 and 4620 shows image parameters corresponding to mean intensity, intensity density, and intensity density standard deviation. In particular, plot 4610 shows the mean intensity of pancreatic cancer samples ranges between values of 100 and 175 and the standard deviations generally above 24. By contrast, plot 4620 shows the mean intensity of non-cancer samples having a range between 50 and 150 and the standard deviation of non-cancerous samples are generally below 24. In some variations, a prediction machine learning model may predict a cancer parameter based on the quantifiable difference in these image parameters.
[0241] Image parameters may be used to distinguish between cancerous and non-cancerous samples. For example, a red mean average image parameter may correspond to an intensity of a red channel across all the pixels in an image. In an eosin-stained sample, a colorectal cancer sample may have a lower red mean average image parameter due to a wider color range in the image data taken of the sample. For instance, in a colorectal cancer sample the colors in the image data may vary from dark purples to light pinks. By contrast, a red mean average image parameter from image data collected from an eosin-stained non-cancerous sample may have more uniform red or magenta tones compared to a cancerous sample. As another example, aAttorney Docket No.: AUMI-004 / 02WO 348385-2089 texture entropy image parameter may correspond to randomness or disorder in the texture of image data taken of a sample. A higher texture entropy image parameter may correspond to a more complex or less uniform texture. A texture entropy image parameter taken from image data of a colorectal cancer sample, for example, may be high due to random, pleomorphic puncta. By contrast, a texture entropy image parameter taken from image data of a non-cancerous sample may be distinguishably low due to a presence of smooth and / or uniform ECMBs. As yet another example, a Gabor energy variance image parameter may correspond to a variability, expressed by different orientations or frequencies, within an image. A higher Gabor energy variance suggests a wider range of dominant texture patterns. As such, a Gabor energy variance image parameter taken from image data of a colorectal cancer sample, for example, may be high due to greater texture and color expressed by the cancerous ECMBs on the microfluidic chip. By comparison, a Gabor energy variance image parameter taken from image data of a non- cancerous sample may be lower due to the ECMBs exhibiting more heterogeneity and / or more homogeneousness.
[0242] The methods described herein may be applied to black and white or monochrome images. For these, color-based parameters may be adapted to a grayscale intensity average. Texture-based parameters, such as texture entropy and Gabor energy variance, may be particularly effective as they rely on the spatial arrangement of pixel intensities, which is a fundamental aspect of image texture regardless of color.
[0243] In some variations, a prediction machine learning model may perform parameter selection to determine one or more predictive image parameters based on the ability of the image parameter to accurately distinguish between cancer parameters. For example, an image parameter may be a better predictor of a first cancer type compared to second cancer type, a first cancer stage compared to a second cancer stage, and / or a cancer parameter based on a first morphological feature compared to the same cancer parameter based on a second morphological feature. In some variations, a prediction machine learning model may selectively emphasize (e.g., weight) an image parameter based on a morphological feature present in the ECMB data. Additionally or alternatively, one or more image parameters may be excluded from analysis by the prediction machine learning model. For example, a prediction machine learning model trained to analyze regions of interests containing a morphological feature corresponding toAttorney Docket No.: AUMI-004 / 02WO 348385-2089 heterogeneity in staining patterns may emphasize image parameters corresponding to colors and textures in the image (e.g., high-frequency energy, low-frequency energy, and edge entropy).
[0244] In some variations, a prediction machine learning model may selectively emphasize (e.g., weight) an image parameter using supervised or unsupervised machine learning techniques. Using unsupervised techniques, the image parameters selectively emphasized may not correspond to a morphological feature but may nevertheless be used in predicting a cancer parameter. For instance, in some variations using unsupervised techniques, k-means clustering may be applied to a plurality of image parameters to identify an image parameter or group of image parameters that when considered together best distinguish between cancerous and non- cancerous samples. In some variations, principal component analysis (PCA) may be applied to one or more predictive image parameters. In some variations, the prediction of the machine learning model may be based on principal components generated in PCA rather than or in addition to the image parameters. A principal component may be selectively emphasized in the same or similar manner that an image parameter is selectively emphasized (e.g., using K-means clustering).
[0245] In some variations, a prediction machine learning model may be trained to distinguish between cancerous samples and non-cancerous samples based on unlabeled data by analyzing image parameters and / or principal components of the images parameters, generated using principal component analysis (PCA), thereby enabling more data to be used for validation and testing. In some variations, a prediction machine learning model may predict a cancer parameter using a majority voting algorithm amongst one or more of the selected image parameters and / or principal components. A prediction machine learning model may predict a cancer parameter using one or more clustering techniques (e.g., K-means clustering) to separate cancerous samples and non-cancerous samples in unlabeled data.
[0246] In some variations, as mentioned herein, image parameters may be generated and selected to train a prediction machine learning model to predict a cancer parameter based on the image parameters. FIG.92 depicts a flowchart representation of a method 9200 of generating a prediction machine learning model for analyzing image parameters. The method 9200 may include receiving ECMB image data corresponding to ECMBs separated from a biological fluid 9202. For example, ECMB samples corresponding to one or more of stained non-cancerous ECMBs, breast cancer ECMBs, colorectal cancer ECMBs, prostate cancer ECMBs, lung cancerAttorney Docket No.: AUMI-004 / 02WO 348385-2089 ECMBs, pancreatic cancer ECMBs, or pre-cancer ECMBs of any of the preceding cancer types may be received. The ECMBs may have been processed on a microfluidic chip as described herein. The ECMBs may be imaged on the microfluidic chip and the ECMB images may be preprocessed as described herein.
[0247] In some variations, the method 9200 may optionally include applying a histogram to the ECMB image data 9204. The histogram may improve the visibility of the ECMBs processed on the chip in the ECMB image data. The histogram may adjust one or more of an exposure, highlights, shadows, whites, blacks, brightness, contrast and / or a color channel of an ECMB image data to ensure that subtle morphological features of ECMBs, which might otherwise be obscured by variations in lighting or sample preparation, are clearly discernible for analysis (e.g., analysis by a prediction machine learning model). The histogram may be based on a type of cancer and / or stain applied to ECMBs.
[0248] ECMB image data may be segmented into one or more regions of interest (ROIs) for image parameter generation and / or selection and / or for analysis by a prediction machine learning model. In some variations, the method 9200 may include segmenting received ECMB image data into one or more regions of interest. The ECMB image data may be segmented by any methods of segmenting described herein. For example, the ECMB image data may be segmented into regions of interest by applying one or more thresholds based on the processing of the ECMBs (e.g., processing using eosin stain, processing using IF stains). In some variations, ECMB image data may be segmented into regions of interest by applying one or more thresholds associated with one or more of a hue, a saturation, a brightness, and / or size of the image. For instance, ROIs of ECMB image data of ECMBs processed with an eosin stain may be segmented with a hue range of about 206 to about 255, a saturation range of about 43 to about 255, and a brightness range of about 74 to about 255, including all subranges and values therein, with a size threshold of about 1000, for noise reduction. Similarly, ROIs of ECMB image data of ECMBs processed with one or more IF stains may be segmented with a hue range of about 0 to about 255, a saturation range of about 0 to about 255, and a brightness range of about 42 to about 255, including all subranges and values therein, with a size threshold of about 100, for noise reduction.
[0249] Image parameters may be extracted from ECMB image data and / or one or more ROIs associated with the ECMB image data. The method 9200 may include generating imageAttorney Docket No.: AUMI-004 / 02WO 348385-2089 parameters from one or more ROIs 9208. The image parameters may be generated after the biological fluid containing the ECMBs has been processed on the microfluidic chip, the processed ECMBs have been imaged, the ECMB image data has been received, the ECMB image data has been preprocessed, and / or the ECMB image data has been segmented into one or more ROIs. A plurality of image parameters may be generated from each of the one or more regions of interest. The image parameters may include one or more of any of the image parameters described herein. Generally image parameters may be associated with a color, a texture, an edge, a wavelet, and / or the like of the one or more ROIs.
[0250] In some variations, the plurality of image parameters may include about 10 or more image parameters, about 50 or more image parameters, about 100 or more image parameters, about 200 or more image parameters, about 400 or more image parameters, about 600 or more image parameters, about 800 or more image parameters, or about 1000 or more image parameters, including all subranges and values therein. For example the plurality of image parameters may be about 5 image parameters, about 10 image parameters, about 20 image parameters, about 30 image parameters, about 40 image parameters, about 50 image parameters, about 60 image parameters, about 70 image parameters, about 80 image parameters, about 90 image parameters, about 100 image parameters, about 150 image parameters, about 200 image parameters, about 250 image parameters, about 300 image parameters, about 400 image parameters, about 500 image parameters, about 600 image parameters, about 800 image parameters, or about 1000 image parameters. In some variations, each image parameter of the plurality of image parameters may be averaged based on the value of that image parameter associated with a respective ROI.
[0251] In some variations, the method 9200 may optionally include normalizing the image parameters 9210. For example, the plurality of image parameters may be normalized (e.g., min- max scaled) such that all image parameters are rescaled to range from 0 to 1. In some variations, the method 9200 may optionally include dividing the ECMB data, including the image parameters generated for the one or more ROIs, into a training set and test set. The training set may be used to select a subset of image parameters and / or generate a prediction machine learning model to predict a cancer parameter based on the image parameters. The test set may be used to evaluate the prediction machine learning model. The training set and test set may be mutually exclusive, such that no ECMB data of the same sample appears in both sets. In someAttorney Docket No.: AUMI-004 / 02WO 348385-2089 variations, the ECMB data may not be divided. For example, the entire data set may be used to train and / or cross validate the prediction machine learning model to analyze image parameters.
[0252] In some variations, less than the entire plurality of generated image parameters may be used to train a prediction machine learning model and / or used by the model to predict a cancer parameter. The method 9200 may include selecting a subset of image parameters 9214. Selecting a subset of image parameters may reduce the dimensionality of inputs to the prediction machine learning model and, thus, improve the efficiency of the model. The subset of image parameters may be selected based on the architecture of the machine learning model (e.g., logistic regression, XGBoost), the type of cancer to be predicted, the stage of cancer to be predicted, and / or the processing of the imaged ECMBs (e.g., histological processing, IF processing), or combinations thereof. In some variations, the subset of image parameters may be determined by eliminating image parameters that vary little between cancerous and noncancerous samples and thus are of less importance to a prediction machine learning model. The subset of image parameters may be selected using one or more feature selection techniques including Recursive Feature Elimination (RFE), Boruta, Mutual Information (MI), and SHapley Additive exPlanations (SHAP).
[0253] In some variations, the number of features selected may be selected based on the estimated performance the prediction machine learning model. The subset of image parameters may include about 5 to about 300 image parameters, about 5 to about 200 image parameters, about 5 to about 100 image parameters, about 5 to about 50 image parameters, about 20 to about 200 image parameters, about 20 to about 100 image parameters, about 20 to about 80 image parameters, about 60 to about 200 image parameters, about 60 to about 100 image parameters, about 80 to about 200 image parameters, about 80 to about 120 image parameters, about 80 to about 100 image parameters, or about 100 image parameters to about 200 image parameters, including all subranges and values therein. For example the subset of image parameters may be about 5 image parameters, about 10 image parameters, about 20 image parameters, about 30 image parameters, about 40 image parameters, about 50 image parameters, about 60 image parameters, about 70 image parameters, about 80 image parameters, about 90 image parameters, about 100 image parameters, about 150 image parameters, or about 200 image parameters.
[0254] A prediction machine learning model may be trained to predict a cancer parameter based on a selected subset of image parameters. In some variations, the method 9200 mayAttorney Docket No.: AUMI-004 / 02WO 348385-2089 include training and / or cross validating a prediction machine learning model 9216. In some variations, the prediction machine learning model may be trained on the training set. The ECMB data including the plurality of image parameters may be labeled with one or more cancer parameter predictions during training. For example, training may comprise the prediction machine learning model receiving as input a selected subset of image parameters. In some variations, training and cross validating the machine learning model may comprise selecting an architecture for the model. The architecture may include one or more classifiers including Random Forest (RF), Logistic Regression (LogReg), Support Vector Machine (SVM), XGBoost, LightGBM, CatBoost, Gradient Boosting Machine (GBM), K-Nearest Neighbors (KNN), Naive Bayes, Decision Tree, Linear Discriminant Analysis (LDA), Quadratic Discriminant Analysis (QDA), Stochastic Gradient Descent Classifier (SGDClassifier), Voting Classifier, Stacking Classifier, Bagging Classifier, Explainable Boosting Machine (EBM), and the like or combinations thereof. The architecture may be selected by training a plurality of models comprising different architectures and cross validating the models to estimate the performance of each model architecture. For example, the prediction machine learning model may be trained using 5-fold cross-validation to estimate the performance of a plurality of combinations of image parameter selection methods and prediction machine learning architectures. In some variations, the trained prediction machine learning model may be stored to analyze additional ECMB data including image parameters.
[0255] The method 9200 may optionally further comprise testing the trained prediction machine learning model 9218. The model may be tested on the testing set comprising ECMB data separated from the ECMB data used to train the model. Similar to how the model may be trained, a prediction machine learning model may receive as inputs only a selected subset of image parameters for testing. Testing and / or cross validating the prediction machine learning model may comprise evaluating the model based on one or more metrics to quantify the performance of the model. These metrics may include one or more of an area under the receiver operating characteristic (ROC) curve (AUROC) to quantify how well the model separates cancerous samples from controls (1.0 being perfect separation and 0.5 being random guessing), sensitivity, specificity, PPV, NPV, and Youden’s J. In some variations, testing the prediction machine learning model may comprise optimizing a threshold of the model. For example, a threshold (e.g., classification threshold) may be selected to maximize Youden’s J to produce a prediction machine learning model with balanced sensitivity and specificity performance.Attorney Docket No.: AUMI-004 / 02WO 348385-2089 Scoring Rubrics
[0256] In some variations, cancer prediction based on scoring rubrics as described herein may be used for diagnosis, prognosis, detecting minimal residual disease, and / or treatment of a subject. For example, a cancer parameter may be predicted with high sensitivity and specificity based on human analysis of ECMB morphology of stained plasma ECMBs. FIG.11A is a table of a scoring rubric 1100 of a method of predicting colorectal cancer (CRC) as either a non- cancerous ECMB sample or a colorectal cancer ECMB sample based on five structure classes. Higher weight was given to morphological feature sets having higher presence in cancer ECMBs and lower presence in non-cancerous ECMBs. The scores from each weight may be summed to give a total score corresponding to a cancer prediction as shown in table 1110. For example, an ECMB sample that does not have any of the structures A-E corresponds to a total score of 0 and may be predicted as a definite control (e.g., non-cancerous). Conversely, if a sample matches each cancer descriptor for classes A-E, then the sample would have a total score of 7 corresponding to a definite colorectal cancer tumor (e.g., positive for colorectal cancer).
[0257] FIG.11B are images 1120 of exemplary scoring for an eosin-stained non-cancerous ECMB sample on a microfluidic chip based on the scoring rubric 1100. FIG.11C are images 1130 of exemplary scoring for an eosin-stained colorectal cancer ECMB sample on a microfluidic chip based on the scoring rubric 1100. In some variations, the cancer predictions based on the human analysis of ECMB morphology using the scoring rubrics described herein may be used to generate training data for training a prediction machine learning model.
[0258] Based on logistic regression analysis, for each additional class structure identified in the ECMB sample (e.g., higher total score), the probability of tumor presence increased by 77%. Accordingly, the number of cancerous class structures identified may be predictive of a cancer prediction. In one example, cancer prediction based on the scoring rubric 1100 has a sensitivity of 71.43%, a specificity of 83.33%, a PPV of 75.00%, a NPV of 80.65%, and a correct classification rate of 78.43%. Furthermore, the area under the Receiver Operating Characteristic (ROC) curve (AUC) is 0.8032 (95% CI: 0.677-0.929). AUC measures the model's ability to discriminate between classes, with higher values indicating better discrimination. In another example, cancer prediction based on the scoring rubric 1100 was tested 25 samples including 10 colorectal cancer samples and 15 control samples. The cancer prediction based on the scoring rubric 1100 had an accuracy of 88%, a sensitivity of 80%, a specificity of 93%, a PPV of 89%, aAttorney Docket No.: AUMI-004 / 02WO 348385-2089 NPV of 88%. Furthermore, the cancer prediction generated from the scoring rubric 1100 had an AUC of 0.87 demonstrating a strong ability discriminate between colorectal cancer and non- cancer using this method.
[0259] In some variations, prostate cancer prediction may be based on a scoring rubric and used for diagnosis, prognosis, and / or treatment of a subject. For example, pancreatic cancer (e.g., PDAC) may be predicted with high sensitivity and specificity based on human analysis of ECMB morphology of stained plasma ECMBs. FIG.27A is a table 2700 of a scoring rubric of a method of predicting pancreatic cancer as one of a non-cancerous ECMB sample (e.g., control), a possible pancreatic cancer ECMB sample (e.g., equivocal PDAC), or a pancreatic cancer ECMB (e.g., PDAC) based on morphological features. For example, FIG.27B includes images 2702-2712 corresponding to PDAC ECMB samples 2703, 2705, 2707, 2709, 2711, 2713 where the most prominent feature set of PDAC is the presence of scattered or clustered large, basophilic to amphophilic sharply demarcated, pebble-like, opaque granules, and less frequently eosinophilic bubble-like, slightly translucent vesicles, easily visualized at between around 20% to around 30% magnification. These large granules / vesicles tend to overlay the larger flakes. Less frequently, these large granules / vesicles are visualized without associated background flakes. The larger flakes in the background of PDAC had densely eosinophilic, paste-like, smudgy, granular, or bubbly consistency. FIGS.27C and 27D includes images 2714-2724 corresponding to equivocal PDAC samples 2715, 2717, 2719, 2721, 2723, 2725, 2727, 2729 where characteristic granules / vesicles are present without flakes having basophilic / amphophilic large granules / vesicles or when flakes are present without free-floating basophilic / amphophilic large granules / vesicles. For example, image 2722 depicts smooth veil-like flake 2723 and crystalline-like angulated granule 2725. Image 2724 depicts crystalline-like flake 2727 with overlying poorly defined eosinophilic granules 2729.
[0260] FIG.27E includes images 2726-2730 corresponding to control ECMB samples 2731, 2733, 2735 having either translucent gray, irregular, crystalline-like granules or small eosinophilic granules not easily visible at between about 20% and about 30% magnification. The background flakes tended to be smooth, translucent, cellophane-like, veil-like, and / or angulated crystalline-like. In some variations, angulated folds in the flakes could mimic PDAC basophilic granules where the smooth circular sharply circumscribed edge served to separate the PDAC from folding mimicker. For example, image 2726 includes small basophilic crystalline-likeAttorney Docket No.: AUMI-004 / 02WO 348385-2089 angulated granules 2737 overlaying a smooth veil-like flake 2731, image 2728 includes smooth crystalline-like flakes 2733 with focal irregularities that mimic pebble-like deposits of PDAC and which overlay poorly defined eosinophilic granules, and image 2730 includes irregular angulated crystalline-like flakes 2735 with folds that mimic basophilic PDAC granule / vesicle.
[0261] FIGS.27F-27H include images 2732-2740 of false positive control ECMB samples 2739, 2741, 2743. For example, images 2732-2736 include intense eosinophilia that may be mistaken for granules and flakes corresponding to a PDAC ECMB sample. FIGS.27G and 27H are exemplary images 2738, 2740 of false positive control ECMB samples 2739, 2741 having an intensely eosinophilic background in the right-hand column. FIG.27I corresponds to an image 2742 of a background of a control ECMB sample 2743. FIG.27J corresponds to an image 2744 of a background of a PDAC ECMB sample 2745. FIG.27K corresponds to an image 2746 of a frequent PDAC background of a PDAC ECMB sample 2747 having less intense eosinophilia when compared to the false-positive control ECMB samples.
[0262] Possible exclusion criteria (may be grouped with equivocal PDAC to increase sensitivity or with equivocal controls to increase specificity) may include intense, deeply eosinophilic staining in the 4 um gap region of the chip that was found to result in a high rate of false positives. Intensely eosinophilic granules surrounding each circular column in the large gap region of the chip may be associated with a false positive.
[0263] FIG.27L is a binary classification table corresponding to an illustrative variation of a scoring rubric method of predicting pancreatic cancer. Table 2748 demonstrates that analysis of ECMB morphology provides high accuracy and precision in differentiating between healthy tissue and pancreatic cancer. For example, a pathologist identified feature sets consistent with pancreatic cancer in early and late-stage disease. The pathologist identified feature sets are consistent with healthy controls and developed a scoring method to identify the morphology of healthy control plasma as described herein. Forty images were used for training (e.g., 28 controls, 12 PDAC) and 29 images used for validation and testing (e.g., 28 controls, 12 PDAC).
[0264] FIG.53 is a table 5300 of a scoring rubric of a method of predicting prostate cancer based on morphological feature sets as one of a non-cancerous ECMB sample (e.g., definite healthy control), a probable healthy control ECMB sample, a possible healthy control ECMB sample, a possible prostate cancer ECMB sample, a probable prostate cancer ECMB, or aAttorney Docket No.: AUMI-004 / 02WO 348385-2089 prostate cancer ECMB sample. Higher weight may be given to morphological feature sets having higher abundance (e.g., presence) in prostate cancer ECMBs and lower abundance in non-cancerous ECMBs. For example, a first and second morphological feature (e.g., deposit complexes A, B) had the highest determinative power in predicting prostate cancer and are assigned a score of 2 points. A third, fourth, and fifth morphological feature (e.g., deposit complexes C, D, E) where partially determinative of prostate cancer but less determinative than the first and second feature and are assigned a score of 1 point. No weight was given to morphological feature sets (e.g., healthy deposit complexes A-E) characteristic of non-cancerous controls samples. The scores from each weight may be summed to give a total score corresponding to a cancer prediction. For example, an ECMB sample that does not have any of the structures A-E corresponds to a total score of 0 and may be predicted as a definite healthy control. Conversely, if a sample matches each cancer descriptor for deposit complexes A-E, then the sample would have a total score of 7 corresponding to a cancer prediction of a definite prostate cancer tumor.
[0265] FIG.54 includes images 5402-5416 corresponding to prostate cancer ECMB samples identified as prostate cancer where one or more of deposit complexes A-E corresponding to prostate cancer are present. For example, image 5402 includes an opaque ECMB, molded to the pillar like paste with a curvilinear edge and was identified as exhibiting deposit complex B, thereby scoring 2 points. Image 5402 also includes uniform and homogeneous staining of the ECMB and was identified as exhibiting deposit complex C, thereby scoring an additional point. Image 5402 also includes thick amorphous bands, greater than the distance between adjacent diagonal columns in the same row and was identified as exhibiting deposit complex D, thereby scoring an additional point. Image 5402 finally includes an ECMB encircle more than 180 degrees of a pillar and was identified as exhibiting deposit complex E, thereby scoring an additional point. Accordingly, image 5402 was given a total score of 5, indicating a probable prostate cancer sample. Images 5404, 5406, 5412, and 5416 exhibit deposit complexes B, C, D, and E and were given a total score of 5 indicating probable prostate cancer.
[0266] Next, image 5408 includes dense eosinophilic smudgy granules that overly coat and encircle entire pillars and was identified as exhibiting complex A, thereby scoring 2 points. Image 5408 also includes deposit complex B scoring an additional 2 points, deposit complex C scoring an additional point, and deposit complex E scoring an additional point, and none of theAttorney Docket No.: AUMI-004 / 02WO 348385-2089 healthy deposit complexes. Thus, image 5408 was given a total score of 6, indicating a probable prostate cancer sample. Image 5410 was identified as exhibiting deposit complexes B, C and D for a total of 4 points, indicating possible prostate cancer. Image 5414 was identified as exhibiting deposit complexes A, B, C, and E for a total score of 6 points, indicating probable prostate cancer.
[0267] FIG.55 is a binary classification table corresponding to a scoring rubric method of predicting prostate cancer. Table 5500 demonstrates that analysis of ECMB morphology provides high accuracy and precision in differentiating between healthy tissue and prostate cancer. For example, a pathologist identified features consistent with prostate cancer in early and late-stage disease. The pathologist identified features consistent with healthy controls and developed a scoring method to identify the morphology of healthy control plasma. For example, a total score of 4 or above was found to best classify between healthy controls and prostate cancer.86 images were used for training (e.g., 55 controls, 31 prostate cancer) and 18 images were used for validation and testing.
[0268] In some variations, breast cancer prediction may be based on a scoring rubric and used for diagnosis, prognosis, detecting minimal residual disease, and / or treatment of a subject. For example, breast cancer may be predicted with high sensitivity and specificity based on human analysis of ECMB morphology of stained plasma ECMBs. Breast cancer ranks as the second leading cause of death in women. However, screening for breast cancer at an early stage may increase the chances for survival. Conventional screening methods include mammography, MRI, and direct pathological biopsy. However, conventional breast cancer screening has limitations, such as the variability of healthy tissue in X-rays, lower sensitivity in younger patients, and dense breast tissue, which can make it difficult to spot cancers. Generally, screening mammograms are less effective in younger individuals due to their dense breast tissue. Sometimes, mammography fails to detect certain breast cancers due to the cancer's location or the density of the breast tissue. For example, about 25% of cancers in women age 40-49 are not detected by a screening mammogram. On average, 9% of people screened through BC Cancer Breast Screening will require additional testing to examine a specific area of the breast more closely. Additionally, 95% of people recalled for further testing do not have cancer.
[0269] The method of breast cancer prediction described herein may predict between non- cancerous tissue and breast cancer tissue based on a scoring rubric and analysis of morphologicalAttorney Docket No.: AUMI-004 / 02WO 348385-2089 feature sets derived from stained (e.g., eosin) ECMBs isolated on a microfluidic chip. For example, an ECMB sample having any one or more characteristic morphological feature sets (e.g., score of 1-3) may correspond to a breast cancer ECMB sample. Conversely, the absence of any of the characteristic morphological feature sets (e.g., score of 0) may correspond to a non- cancerous (e.g., control plasma) ECMB sample.
[0270] Generally, a breast cancer ECMB sample may have one or more of the following morphological feature sets: 1) Dense, thick, textured, cohesive bandlike deposit with overlying dense eosinophilic large vesicles / granules (easily seen at 50% magnification). The complex of vesicles / granules and band-like deposit had irregular, granular, pseudopod-like margins; 2) Large eosinophilic granules (easily seen at 50% magnification) in a background of eosinophilic band-like or string-like deposits; and 3) Eosinophilic band-like deposit with feathery or finely linear consistency and at least focally dense consistency (i.e. not entirely translucent). A higher number of these morphological feature sets in an ECMB sample may indicate a higher probability of breast cancer.
[0271] Generally, a non-cancerous ECMB sample may have one more of the following morphological feature sets: 1) Translucent, smooth veil-like, band-like or geometric deposit with no pseudopod-like margin; 2) Small eosinophilic or gray (any size) granules; 3) Large eosinophilic granules without an associated string-like or band-like deposit; and 4) Finely translucent deposit with indistinct borders and feathery-like consistency but without any dense consistency.
[0272] FIGS.28A and 28B include images 2800-2806 corresponding to breast cancer ECMB samples having a first morphological feature 2803, 2805 corresponding to dense, thick, textured, cohesive bandlike deposit with overlying dense eosinophilic large vesicles / granules (easily seen at 50% magnification). The complex of vesicles / granules and band-like deposits may have irregular, granular, pseudopod-like margins. FIG.28C includes images 2808-2812 corresponding to breast cancer ECMB samples having a second morphological feature 2807 corresponding to large eosinophilic granules or vesicles (easily seen at 50% magnification) in a background of eosinophilic band-like or string-like deposits. FIG.28D includes images 2814- 2818 corresponding to breast cancer ECMB samples having a third morphological feature 2809, 2811, 2813 corresponding to eosinophilic band-like deposit with feathery or finely linear consistency and at least focally dense consistency (i.e., not entirely translucent).Attorney Docket No.: AUMI-004 / 02WO 348385-2089
[0273] FIG.28E includes images 2820-2826 corresponding to control ECMB samples having a fourth morphological feature 2815, 2817, 2819, 2821 having translucent, smooth veil-like, band-like or geometric deposits with no pseudopod-like margin, small granules, not easily seen at 50% magnification. FIG.28F includes images 2828 and 2830 corresponding to control ECMB samples having a fifth morphological feature 2823 having gray (not eosinophilic) granules in a background of string-like deposits (e.g., image 2828) or gray and eosinophilic large granules / vesicles 2825 without appreciable string-like or band-like deposit in the background (e.g., image 2830). FIG.28G includes image 2832 corresponding to a control ECMB sample having a sixth morphological feature 2827 having finely translucent deposits with indistinct borders and feathery-like consistency but without any dense consistency, and large gray (not eosinophilic) granules / vesicles. Images 2800-2832 of FIGS.28A-28G are each shown at 150% magnification. FIG.28H is a binary classification table corresponding to a scoring rubric method of predicting breast cancer. Table 2834 demonstrates that analysis of ECMB morphology provides high accuracy and precision in differentiating between healthy tissue and breast cancer.
[0274] In some variations, lung cancer prediction may be based on a scoring rubric and used for diagnosis, prognosis, and / or treatment of a subject. For example, lung cancer may be predicted with high sensitivity and specificity based on human analysis of ECMB morphology of stained plasma ECMBs. In some variations, as described with respect to FIG.50A and 51A, a cancer parameter may be predicted based on one or more scoring rubrics. FIG.50A and 51A are tables 5000, 5100 of a scoring rubric of a method of predicting lung cancer based on morphological feature sets as one of a non-cancerous ECMB sample (e.g., healthy control), a probable healthy control ECMB sample, a possible healthy control ECMB sample, a possible lung cancer ECMB sample, a probable lung cancer ECMB sample, and a lung cancer ECMB sample. In some variation, an ECMB sample having any one or more morphological feature sets characteristic of a lung cancer ECMB sample may score a predetermined number of points (e.g., score of 1-3 points). Conversely, the absence of any of the characteristic morphological feature sets (e.g., score of 0) may correspond to a non-cancerous (e.g., control plasma) ECMB sample. In some variations, an ECMB sample having one or more morphological feature sets characteristic of both a healthy control and lung cancer (e.g., overly intense staining, grey deposits with no staining) may be excluded from analysis to avoid false positives or false negatives.Attorney Docket No.: AUMI-004 / 02WO 348385-2089
[0275] In some variations, higher weight may be given to morphological feature sets having higher abundance in cancer ECMBs and lower abundance in non-cancerous ECMBs. For example, a first morphological feature (e.g., deposit complex 1) may have the highest determinative power in predicting lung cancer and is assigned a score of 3 points in table 5000. FIG.50B includes images 5012-5016 corresponding to lung cancer ECMB samples identified as having a first morphologic feature characteristic of lung cancer (e.g., deposit complex 1). Images 5012-5016 include ECMBs characterized by one or more of a dense, textured, paste-like, crumb- like or band-like lightly eosinophilic deposit with overlying large (e.g., easily visible at 30% magnification) eosinophilic-to-basophilic granules or vacuoles. A second morphological feature characteristic of lung cancer (e.g., deposit complex 3) may have determinative power in predicting lung cancer and is assigned a score of 2 points in table 5000. FIG.50D includes images 5032 and 5034 corresponding to lung cancer ECMB samples identified as having a second morphologic feature characteristic of lung cancer (e.g., deposit complex 3). Images 5032 and 5034 include ECMB characterized by one or more of a dense, textured, paste-like, crumb- like or band-like lightly eosinophilic deposit without overlying large (e.g., easily visible at 30% magnification) eosinophilic-to-basophilic granules or vacuoles. A third morphological feature (e.g., deposit complex 2) may have lower determinative power in predicting lung cancer and is assigned a score of 1 point in table 5000. FIG.50C includes images 5022 and 5024 corresponding to lung cancer ECMB samples identified as having the third morphologic feature characteristic of lung cancer (e.g., deposit complex 2). Images 5022 and 5024 include ECMB characterized by large eosinophilic or basophilic round or ovoid granules and vacuoles (easily visible at 30% magnification) in a background of string-like deposits in a well stained preparation (e.g., pearls on a string-like appearance).
[0276] In some variations, an ECMB sample having one or more excluded morphological feature sets characteristic of both a healthy control and lung cancer including overly intense staining or grey deposits with no staining may be erroneous and excluded from analysis. For example, FIG.50E includes images 5042 and 5044 corresponding to erroneous ECMB samples identified as having a first excluded morphological feature (e.g., exclusion criteria 1). In particular, images 5042 and 5044 include ECMBs characterized by intense eosinophilic staining, thereby rendering analysis of the morphology to be difficult. FIG.50F includes images 5052 and 5054 corresponding to ECMB samples identified as having a second excluded morphological feature (e.g., exclusion criteria 2). For example, images 5052 and 5053 include ECMBAttorney Docket No.: AUMI-004 / 02WO 348385-2089 characterized by edge precipitation that mimics the first morphological feature corresponding to lung cancer.
[0277] In some variations, no weight may be given to morphological feature sets characteristic of non-cancerous controls samples. For example, FIG.51B includes images 5110-5130 of ECMB morphologies corresponding to non-cancerous morphological feature sets (e.g., healthy control complexes 1-3) and are assigned a score of 0 points on table 5100. Image 5110 includes ECMB corresponding to a first non-cancerous morphological feature (e.g., healthy control deposit 1) characterized by one or more of a sharply demarcated, angulated, or geometric deposit that does not otherwise fit the characterization of the morphological feature sets corresponding to lung cancer. Image 5120 includes ECMB corresponding to a second non-cancerous morphological feature (e.g., healthy control deposit 2) characterized by a translucent veil-like deposit with overlying basophilic that does not otherwise fit the characterization of the morphological feature sets corresponding to lung cancer. Image 5130 includes ECMB corresponding to a third non-cancerous morphological feature (e.g., healthy control deposit 3) characterized by translucent veil-like deposit with few small eosinophilic granules that does not otherwise fit the characterization of the morphological feature sets corresponding to lung cancer.
[0278] In some variations, the scores (e.g., point value) from each morphological feature may be summed to give a total point value corresponding to a cancer prediction. For example, an ECMB sample that does not have any of the deposit complexes 1-3 of table 5000 or corresponds to a total score of 0 may be predicted as a definite non-cancerous control. Conversely, if a sample matches each morphologic feature characteristic of lung cancer (e.g., deposit complexes 1-3), then the sample would have a total score of 6 corresponding to a prediction of lung cancer.
[0279] FIG.52 is a binary classification table corresponding to an illustrative variation of a scoring rubric method of predicting lung cancer. Table 5200 demonstrates that analysis of ECMB morphology provides high accuracy and precision in differentiating between non- cancerous tissue and lung cancer. In particular, a pathologist identified features consistent with non-cancerous controls and developed a scoring method to identify the morphology of non- cancerous control plasma. A total score of 3 or above was found to best classify between non- cancerous controls and lung cancer.27 annotated images were used for training (e.g., 10 controls, 17 lung cancer), 35 images were use for validation, and 19 unannotated images used for testing.Attorney Docket No.: AUMI-004 / 02WO 348385-2089 Interpretability of Machine Learning Models and Correlation to Human Methods
[0280] In some variations, the analysis of a prediction machine learning model may be compared to human analysis to increase the interpretability of the prediction machine learning model. Both a prediction machine learning model and a human may classify (e.g., predict a cancer parameter of) a fluid sample based on analysis of one or more regions of interest. The one or more regions of interest may be informative as to which areas or features of an image of the fluid sample each mode of analysis (e.g., prediction machine learning model analysis, human analysis, morphological feature analysis) considers important to its classification of the fluid sample. By comparing the regions of interest identified by the prediction machine learning model and the human analysis, one or more of the explainability, interpretability, and accuracy of the prediction machine learning model may be evaluated. For example, if the machine learning model identifies regions of interest relevant (e.g., important) to its classification of the fluid sample that are the same (e.g., overlap) as regions of interest identified as relevant (e.g., important) in human analysis (e.g., pathologist analysis, trained scientist analysis), then the prediction machine learning model may be interpreting the sample similar to how a trained human would analyze the sample. This correlation between modes of analysis may validate the performance of the prediction machine learning model in predicting fluid samples as cancerous or non-cancerous and increase stakeholder confidence and trust in the prediction and / or prediction machine learning model by demonstrating that the model engages in analysis similar to a human.
[0281] In some variations, a human (e.g., trained scientist, pathologist) may analyze ECMB data according to any of the methods described herein and identify one or more regions of interest for comparison against the analysis of a prediction machine learning model. For example, a trained scientist may analyze a fluid sample based on one or more morphological features as described herein. A human may receive image data of processed ECMBs (e.g., matricles) and first holistically analyze the image data (e.g., analyze at low magnification). The human may consider the abundance, distribution, and morphological features of the processed ECMBs of the image and select one or more possible regions of interest. The human may further evaluate (e.g., analyze at higher magnification) and confirm the possible regions of interest through more detailed morphological feature analysis. In some variations, one or more regions of interest (e.g., confirmed regions of interest) may be identified in the image data by defining aAttorney Docket No.: AUMI-004 / 02WO 348385-2089 bounding box including portion(s) of the image corresponding to the region of interest, as shown in FIGS.71-75. For example, bounding boxes of the human identified regions of interest may be defined manually. In some variations, an importance or relevance score is assigned to each region of interest for comparison to the importance placed on regions of interest identified by a prediction machine learning model. For example, regions of interest including morphological features associated with a positive cancer prediction may be assigned a higher cancer-like relevance score. Regions of interest including morphological features associated with a non- cancer prediction may be assigned a higher control-like relevance score. In some variations, the fluid sample may be predicted as cancerous or non-cancerous based on analysis of the regions of interest.
[0282] In some variations, a cancer parameter prediction by a prediction machine learning model may be analyzed to identify one or more regions of interest for comparison to a cancer parameter prediction by human analysis. A prediction machine learning model may be trained to predict a cancer parameter according to the methods described herein. The prediction machine learning model may analyze the same fluid sample images analyzed by the human. In some variations, the cancer parameter prediction of the prediction machine learning model and the prediction machine learning model itself may be analyzed using a saliency mapping technique. For example, the saliency mapping technique may identify regions of higher relevance in the received image data (e.g., region of interest) to the prediction machine leaning model in forming the cancer parameter prediction. For example, one or more regions of interest may be identified by analyzing the prediction machine learning model using a method of explanation with ranked area integrals (XRAI) and examining the area designated as most critical (e.g., important) to the confidence of the prediction machine learning model. In some variations, one or more regions of interest (e.g., regions of higher relevance to model confidence) identified by the salient mapping technique may be defined by a bounding box, as shown in FIGS.71-75. In some variations, bounding boxes of the prediction machine learning model identified regions of interest may be defined manually based on the analysis of the salient mapping technique.
[0283] In some variations, regions of interest identified in the same fluid sample by human analysis and by a prediction machine learning model may be compared against each other to determine the correlation between the prediction machine learning model analysis and human analysis. Comparing the regions of interest identified by human analysis and the regions ofAttorney Docket No.: AUMI-004 / 02WO 348385-2089 interest identified using the prediction machine learning model may increase stakeholder confidence in the prediction machine learning model by demonstrating that the model at least partially follows scientifically recognized (e.g., validated) analysis. For example, comparing the regions of interest may demonstrate the prediction machine learning model considers the same ECMB morphological features important to its prediction of a fluid sample as a trained human, scientist, and / or pathologist.
[0284] FIG.71 are images 7100a and 7100b of an eosin-stained cancerous fluid sample corresponding to a patient diagnosed with stage I pancreatic cancer (e.g., PDAC I). The fluid sample was analyzed by a trained human and a prediction machine learning model. The trained human and the prediction machine learning model each correctly predicted the sample as PDAC I. Image 7100a shows the regions of interest 7102a, 7104a identified by human analysis of the fluid sample based on morphological features. Region of interest 7102a was identified by the trained human as being of higher relevance to their prediction of the sample as PDAC. Regions of interest 7104a was identified by the trained human as being of secondary relevance to their prediction of the sample as PDAC I. Image 7100b shows the regions of interest 7102b, 7104b identified by salient mapping and manual bounding of the prediction machine learning model analysis of the fluid sample image. Region of interest 7102b was identified as being of higher relevance to the prediction of the sample by prediction machine learning model as PDAC I. Regions of interest 7104b were identified as being of secondary relevance to the prediction of the sample by the prediction machine learning model as PDAC I. The trained human and the prediction machine learning model identified similar higher relevance regions of interest demonstrating a correlation between the two methods of analysis and that the prediction machine learning model analyzes ECMB data based on explainable and interpretable features present in the ECMB data.
[0285] FIG.72 are images 7200a and 7200b of an eosin-stained cancerous fluid sample corresponding to a patient diagnosed with stage III pancreatic cancer (e.g., PDAC III). The fluid sample was analyzed by a trained human and a prediction machine learning model. The trained human and the prediction machine learning model each correctly predicted the sample as PDAC III. Image 7200a shows the regions of interest 7202a identified by human analysis of the fluid sample based on morphological features. Regions of interest 7202a were identified by the trained human as being of higher relevance to their classification of the sample as PDAC III. ImageAttorney Docket No.: AUMI-004 / 02WO 348385-2089 7200b shows the regions of interest 7202b, 7204b identified by salient mapping and manual bounding of the prediction machine learning model analysis of the fluid sample image. Region of interest 7202b was identified as being of higher relevance to the prediction of the sample by prediction machine learning model as PDAC III. Regions of interest 7204b was identified as being of secondary relevance to the prediction of the sample by the prediction machine learning model as PDAC III. The prediction machine learning model identified three higher relevance regions of interest similar to higher relevance regions of interest identified by the trained human, thereby demonstrating a correlation between the two methods of analysis and that the prediction machine learning model analyzes ECMB data based on explainable and interpretable features present in the ECMB data.
[0286] FIG.73 are images 7300a and 7300b of an eosin-stained cancerous fluid sample corresponding to a patient diagnosed with stage II to stage III pancreatic cancer (e.g., PDAC II- III). The fluid sample was analyzed by a trained human and a prediction machine learning model. The trained human and the prediction machine learning model each correctly predicted the sample as PDAC II-III. Image 7300a shows the regions of interest 7302a, 7304a identified by human analysis of the fluid sample based on morphological features. Region of interest 7302a was identified by the trained human as being of higher relevance to their prediction of the sample as PDAC II-III. Region of interest 7304a was identified by the trained human as being of secondary relevance to their prediction of the sample as PDAC II-III. Image 7300b shows the regions of interest 7302b, 7304b identified by salient mapping and manual bounding of the prediction machine learning model analysis of the fluid sample image. Region of interest 7302b was identified as being of higher relevance to the prediction of the sample by prediction machine learning model as PDAC II-III. Regions of interest 7304b were identified as being of secondary relevance to the prediction of the sample by the prediction machine learning model as PDAC II- III. The prediction machine learning model identified three regions of interest similar to the regions of interest identified by the trained human. Additionally, the three regions of interest identified by the prediction machine learning model corresponding to the regions of interest identified by the trained human were similarly important to the prediction machine learning model and trained human, thereby demonstrating a correlation between the two methods of analysis and that the prediction machine learning model analyzes ECMB data based on explainable and interpretable features present in the ECMB data.Attorney Docket No.: AUMI-004 / 02WO 348385-2089
[0287] FIG.74 are images 7400a and 7400b of an eosin-stained control fluid sample corresponding to a non-cancerous patient. The fluid sample was analyzed by a trained human and a prediction machine learning model. The trained human and the prediction machine learning model each correctly predicted the sample as non-cancerous. Image 7400a shows the regions of interest 7402a, 7404a identified by human analysis of the fluid sample based on morphological features. Regions of interest 7402a were identified by the trained human as being of higher relevance to their prediction of the sample as non-cancerous. Region of interest 7404a was identified by the trained human as being of secondary relevance to their prediction of the sample as non-cancerous. Image 7400b shows the regions of interest 7402b, 7404b identified by salient mapping and manual bounding of the prediction machine learning model analysis of the fluid sample image. Region of interest 7402b was identified as being of higher relevance to the prediction of the sample by prediction machine learning model as non-cancerous. Regions of interest 7404b were identified as being of secondary relevance to the prediction of the sample by the prediction machine learning model as non-cancerous. The prediction machine learning model identified a similar higher relevance region of interest compared with the trained human identified regions of interest, thereby demonstrating a correlation between the two methods of analysis and that the prediction machine learning model analyzes ECMB data based on explainable and interpretable features present in the ECMB data.
[0288] FIG.75 are images 7500a and 7500b of an eosin-stained control fluid sample corresponding to a healthy patient. The fluid sample was analyzed by a trained human and a prediction machine learning model. The trained human and the prediction machine learning model each correctly predicted the sample as non-cancerous. Image 7500a shows the regions of interest 7502a identified by human analysis of the fluid sample based on morphological features. Regions of interest 7502a were identified by the trained human as being of higher relevance to their prediction of the sample as non-cancerous. Image 7500b shows the regions of interest 7502b, 7504b identified by salient mapping and manual bounding of the prediction machine learning model analysis of the fluid sample image. Region of interest 7502b was identified as being of higher relevance to the prediction of the sample by prediction machine learning model as non-cancerous. Regions of interest 7504b were identified as being of secondary relevance to the prediction of the sample by the prediction machine learning model as non-cancerous. The prediction machine learning model identified similar important regions of interest compared with the trained human identified regions of interest, thereby demonstrating a correlationAttorney Docket No.: AUMI-004 / 02WO 348385-2089 between the two methods of analysis and that the prediction machine learning model analyzes ECMB data based on explainable and interpretable features present in the ECMB data.
[0289] Regions of interest identified in the same fluid sample by human analysis and by a prediction machine learning model may be quantitatively compared against each other to determine the correlation between the prediction machine learning model analysis and human analysis. Comparing the regions of interest identified by human analysis and the regions of interest identified using the prediction machine learning model may test the accuracy of the prediction machine learning model in identifying important features in the received image data. Quantitative comparison of the prediction machine learning model analysis with human analysis may further increase stakeholder confidence in the prediction machine learning model by demonstrating that the model at least partially follows scientifically validated analysis.
[0290] In some variations, correlation between the prediction machine leaning model and human analysis may be quantified by overlaying the regions of interest identified by each mode of analysis. For example, K-means binary cluster may be applied to image data corresponding to a salient feature mapping of the cancer parameter prediction of the prediction machine learning model and an image including bounded regions of interest corresponding to human ECMB morphology analysis. The clustering process may generate images including a binary mask that distinguishes areas of higher relevance (e.g., foreground, regions of interest) from areas of lower relevance (e.g., background) for each mode of analysis. The masked images may be preprocessed (e.g., resized, formatted) and linearly blended to generate an overlay representation (e.g., overlay image). In some variations, areas of interaction of the human identified regions of interest and prediction machine learning identified regions of interest may be calculated to generate overlapping regions of interest in the overlay representation. The masked images and overlay representation may then be compared to calculate the accuracy (e.g., precision, recall, F1 score) of the region of interest identification by the prediction machine learning model compared to the human analysis. In some variations, the masked images and overlay representation may be compared automatically.
[0291] FIG.76 are masked images 7600a and 7600b and an overlay image 7600c of an eosin- stained cancerous fluid sample corresponding to a patient diagnosed with stage I pancreatic cancer (e.g., PDAC I). The fluid sample was analyzed by a trained human as shown in masked image 7600a and a prediction machine learning model as shown in masked image 7600b.Attorney Docket No.: AUMI-004 / 02WO 348385-2089 Masked image 7600a shows four regions of interest 7602a identified by human analysis of the fluid sample based on morphological features. Masked image 7600b shows three regions of interest 7602b identified by salient feature mapping of the analysis of the prediction machine learning model. Overlay image 7600c shows overlapping region of interest 7602c identified by calculating the intersection of the prediction machine learning model identified regions of interest and the regions of interest identified by human analysis. Images 7600a, 7600b, and 7600c and their respective regions of interest 7602a, 7602b, and overlapping region of interest 7602c demonstrate a correlation between the prediction machine learning model analysis and the human analysis and that the prediction machine learning model analyzes ECMB data based on explainable and interpretable features present in the ECMB data. For example, the prediction machine learning model accurately identified regions of interest corresponding to a region of interest identified by human analysis with a precision (e.g., match ratio) of 0.25, a recall of 1, and a F1 score of 0.4. F1 score, as used here, measures the overlap between regions predicted by the machine learning model and those identified by human analysis. It is the harmonic mean of precision and recall, with values closer to 1 indicating complete agreement.
[0292] FIG.77 are masked images 7700a and 7700b and an overlay image 7700c of an eosin- stained cancerous fluid sample corresponding to a patient diagnosed with stage III pancreatic cancer (e.g., PDAC III). The fluid sample was analyzed by a trained human as shown in masked image 7700a and a prediction machine learning model as shown in masked image 7700b. Masked image 7700a shows four regions of interest 7702a identified by human analysis of the fluid sample based on morphological features. Masked image 7700b shows four regions of interest 7702b identified by salient feature mapping of the analysis of the prediction machine learning model. Overlay image 7700c shows three overlapping region of interest 7702c identified by calculating the intersection of the prediction machine learning model identified regions of interest and the regions of interest identified by human analysis. Images 7700a, 7700b, and 7700c and their respective regions of interest 7702a, 7702b and overlapping regions of interest 7702c demonstrates a correlation between the prediction machine learning model analysis and the human analysis, and that the prediction machine learning model analyzes ECMB data based on explainable and interpretable features present in the ECMB data. For example, the prediction machine learning model accurately identified regions of interest corresponding to three regions of interest identified by human analysis with a precision of 0.75, a recall of 1, and a F1 score of 0.857. An F1 score of about 0.7 and above may indicate moderateAttorney Docket No.: AUMI-004 / 02WO 348385-2089 to strong agreement between model predictions and human analysis, depending on the specific application and clinical requirements.
[0293] FIG.78 are masked images 7800a and 7800b and an overlay image 7800c of an eosin- stained cancerous fluid sample corresponding to a patient diagnosed with stage II to stage III pancreatic cancer (e.g., PDAC II-III). The fluid sample was analyzed by a trained human as shown in masked image 7800a and a prediction machine learning model as shown in masked image 7800b. Masked image 7800a shows five regions of interest 7802a identified by human analysis of the fluid sample based on morphological features. Masked image 7800b shows four regions of interest 7802b identified by salient feature mapping of the analysis of the prediction machine learning model. Overlay image 7800c shows three overlapping region of interest 7802c identified by calculating the intersection of the prediction machine learning model identified regions of interest and the regions of interest identified by human analysis. Images 7800a, 7800b, and 7800c and their respective regions of interest 7802a, 7802b and overlapping regions of interest 7802c demonstrates a correlation between the prediction machine learning model analysis and the human analysis, and that the prediction machine learning model analyzes ECMB data based on explainable and interpretable features present in the ECMB data. For example, the prediction machine learning model accurately identified regions of interest corresponding to three regions of interest identified by human analysis with a precision of 0.6, a recall of 1, and a F1 score of 0.75.
[0294] FIG.79 are masked images 7900a and 7900b and an overlay image 7900c of an eosin- stained non-cancerous fluid sample corresponding to a non-cancerous patient. The fluid sample was analyzed by a trained human as shown in masked image 7900a and a prediction machine learning model as shown in masked image 7900b. Masked image 7900a shows four regions of interest 7902a identified by human analysis of the fluid sample based on morphological features. Masked image 7900b shows a region of interest 7902b identified by salient feature mapping of the analysis of the prediction machine learning model. Overlay image 7900c shows an overlapping region of interest 7902c identified by calculating the intersection of the prediction machine learning model identified regions of interest and the regions of interest identified by human analysis. Images 7900a, 7900b, and 7900c and their respective regions of interest 7902a, 7902b and overlapping region of interest 7902c demonstrates a correlation between the prediction machine learning model analysis and the human analysis, and that the predictionAttorney Docket No.: AUMI-004 / 02WO 348385-2089 machine learning model analyzes ECMB data based on explainable and interpretable features present in the ECMB data. For example, the prediction machine learning model accurately identified a region of interest corresponding to a region of interest identified by human analysis with a precision of 0.25, a recall of 1, and a F1 score of 0.4.
[0295] FIG.80 are masked images 8000a and 8000b and an overlay image 8000c of an eosin- stained non-cancerous fluid sample corresponding to a non-cancerous patient. The fluid sample was analyzed by a trained human as shown in masked image 8000a and a prediction machine learning model as shown in masked image 8000b. Masked image 8000a shows six regions of interest 8002a identified by human analysis of the fluid sample based on morphological features. Masked image 8000b shows seven regions of interest 8002b identified by salient feature mapping of the analysis of the prediction machine learning model. Overlay image 8000c shows five overlapping region of interest 8002c identified by calculating the intersection of the prediction machine learning model identified regions of interest and the regions of interest identified by human analysis. Images 8000a, 8000b, and 8000c and their respective regions of interest 8002a, 8002b and overlapping regions of interest 8002c demonstrates a correlation between the prediction machine learning model analysis and the human analysis, and that the prediction machine learning model analyzes ECMB data based on explainable and interpretable features present in the ECMB data. For example, the prediction machine learning model accurately identified five region of interest identified by human analysis with a precision of 0.75, a recall of 1, and a F1 score of 0.857.
[0296] FIG.81 is a correlation metrics table 8100 of regions of interest identified by a prediction machine learning model compared to regions of interest identified by human (e.g., trained scientist, pathologist) analysis of the same fluid sample. The prediction machine learning model was trained to distinguish between fluid samples associated with pancreatic cancer (PDAC) and non-cancer (e.g., control samples) based on 41 samples. The prediction machine learning model was K-fold cross-validated on an additional 9 samples to avoid overfitting and selection bias. Table 8100 demonstrates high agreement between the regions of interest identified by the prediction machine learning model and the regions of interest identified by the human analysis. Table 8100 shows an average precision of 0.63, an average recall of 1, and an average F1 score of 0.753. This high similarity between the regions of interest identified by the prediction machine leaning model and human analysis increases stakeholder confidence in theAttorney Docket No.: AUMI-004 / 02WO 348385-2089 model by demonstrating that the prediction machine learning model may be classifying samples based on explainable and interpretable features also relied upon by the trained human during morphological feature analysis.
[0297] In some variations, a Cohen’s Kappa coefficient may be calculated to evaluate the agreement between the regions of interest identified by a prediction machine learning model and the regions of interest identified by human analysis. To calculate the kappa score, total ECMB (e.g., matricles) counts may be obtained, and overlapping regions of interest may be identified in a fluid sample by each mode of analysis. Expected agreement (pe), may include a count of one or more ECMBs present in the fluid sample detectable above a predetermined threshold. Observed agreement (po) may include a count of one or more regions of interest identified by a trained scientist. An observed measurement may be the number of instances in which the prediction machine learning model and the human prediction method agreed on their prediction.
[0298] FIG.82 are images 8200 of a cancerous ECMB fluid sample comparing human identified regions of interest and prediction machine learning model identified regions of interest. Images 8200 are an eosin-stained images of a plasma sample from patient diagnosed with Stage I PDAC, processed on a microfluidic chip. The fluid sample was predicted as PDAC I by both a human prediction method and prediction machine learning model based on one or more regions of interest. Image 8200a shows one or more manually annotated regions of interest identified as highly relevant by a trained human using morphological feature-based prediction. Image 8200b shows a salient feature map of relevant regions of interest used by a prediction machine learning model to predict the sample as PDAC I. Image 8200c is masked image of the manually annotated regions of interest identified by a trained human. Image 8200d is a masked image of relevant regions of interest used by the prediction machine learning model to predict the fluid sample as PDAC I. Region of interest 8202 was identified by both the human prediction method and the prediction machine learning model as highly relevant to prediction of the fluid sample as PDAC I.
[0299] FIG.83 are images 8300 of a non-cancerous ECMB fluid sample comparing human identified regions of interest and prediction machine learning model identified regions of interest. Images 8300 are an eosin-stained images of a plasma sample from a healthy patient, processed on a microfluidic chip. The fluid sample was predicted as non-cancerous (e.g., control) by both human analysis and a prediction machine learning model based one or moreAttorney Docket No.: AUMI-004 / 02WO 348385-2089 regions of interest. Image 8300a shows one or more manually annotated regions of interest identified as highly relevant by a trained human using morphological feature-based prediction. Image 8300b shows a salient feature map of relevant regions of interest used by a prediction machine learning model to predict the sample as non-cancerous. Image 8300c is masked image of the manually annotated regions of interest identified by a trained human. Image 8300d is a masked image of relevant regions of interest used by the prediction machine learning model to predict the fluid sample as non-cancerous. Regions of interest 8302 was identified by both human analysis and the prediction machine learning model as highly relevant to prediction of the fluid sample as non-cancerous.
[0300] FIG.84 is a boxplot 8400 of Cohen’s Kappa scores by patient group (e.g., PDAC, control) comparing one or more regions of interest identified by human analysis (using the methods described herein) and one or more regions of interest identified by a prediction machine learning model. The boxplot 8400 of Cohen’s Kappa scores for three PDAC patient samples and three control patient samples shows a range of 0.24-0.84 Kappa values in the PDAC samples and a range of 0.26-0.66 Kappa values for the control samples. The median Kappa value for PDAC patients is 0.49 and the median Kappa score for control patients is 0.43. FIG.85 is a table 8500 of Cohen’s Kappa scores for six patients (3 PDAC, 3 control) comparing regions of interest identified by a human prediction method and regions of interest identified by a prediction machine learning model. The average Kappa value for PDAC patients is 0.52 and corresponds to moderate agreement. The average Kappa score for control patients is 0.45 and corresponds to moderate agreement. The average Kappa score for all patients is 0.49 and corresponds to moderate agreement. II. Systems and Devices
[0301] Generally, the systems and devices described herein may provide high-throughput separation of extracellular matrix bodies from a biological fluid. The separated ECMBs and remaining biological fluid (or portions or fractions thereof) may then be used for diagnosing, prognosing, or in helping to determine a treatment plan for a subject, and the effectiveness of such treatment plan. A block diagram of an exemplary system 1200 is depicted in FIG.12. The system 1200 may comprise one or more of a holder 1212, a reservoir 1213, a robot 1214, a loader 1215, at least one microfluidic chip 1216, a chip connector 1218, a manifold 1220, a negative pressure source 1222, one or more sensors 1224, an input device 1226, a processorAttorney Docket No.: AUMI-004 / 02WO 348385-2089 1228, a memory 1230, a communication device 1232, and an output device 1234, each of which are described in more detail herein.
[0302] In some variations, the holder 1212 (e.g., material storage) may be configured to store one or more fluids (e.g., biological fluid) for transfer to the microfluidic chip(s) 1216. The holder 1212 may comprise one or more of a tray and container. In some variations, the reservoir 1213 may be configured to receive a predetermined fraction of the biological fluid (e.g., waste fluid) from the microfluidic chip 1216. For example, a non-ECMB fraction of the biological fluid may be transferred to and stored in the reservoir 1213 for further processing (e.g., separation, analysis) and / or disposal. Additionally or alternatively, the system may comprise one or more preprocessing components such as a material separation assembly, aliquoting assembly, a biorepository, and the like.
[0303] In some variations, the robot 1214 may be configured to transfer fluids (e.g., biological fluid, reagent) from holder 1212 to at least one microfluidic chip 1216. In some variations, the microfluidic chip 1216 may be configured to receive and process a biological fluid. For example, the microfluidic chip 1216 may include a channel (e.g., microfluidic channel) and at least one obstruction (e.g., pillar) configured to separate and hold (e.g., trap) a predetermined fraction of a biological fluid within the microfluidic chip 1216 while the remaining fraction flows out of the microfluidic chip 1216. In some variations, the chip connector 1218 may be configured to hold at least one microfluidic chip 1216. The microfluidic chip 1216 may be releasably coupled to the chip connector 1218 to facilitate cleaning and help reduce set-up time. For example, the microfluidic chip 1216 may be a disposable component and the chip connector 1218 may be a durable component. In some variations, the chip connector 1218 may comprise disposable connectors (e.g., inlet connector, outlet connector) configured to facilitate negative pressure applied to the microfluidic chip 1216 by the negative pressure source 1222 in order to perfuse and flow the biological fluid through the microfluid chip 1216.
[0304] In some variations, a loader 1215 may be configured to automatically transfer one or more fluids and disposable components into and out of the system 1200 without human interventions, thus maintaining sterility and facilitating high throughput processing. For example, the loader 1215 may be configured to transfer one or more fluids (e.g., biological fluid, reagent, stain, buffer, waste fluid) and disposable components (e.g., microfluidic chip) to and from one or more of the holder 1212, reservoir 1213, and microfluidic chip 1216. In particular,Attorney Docket No.: AUMI-004 / 02WO 348385-2089 the loader 1215 may be configured to transfer a plurality of blood samples to the holder 1212, a plurality of stains to the reservoir 1213, and a plurality of microfluidic chips 1216 to the chip connectors 1218. Conversely, the loader 1215 may be configured to remove waste fluid from the reservoir 1213, microfluidic chips 1216 from the system 1200, and fluid stored in the holder 1212 and reservoir 1213.
[0305] In some variations, the loader 1215 may comprise one or more conveyor tracks configured to move one or more fluids and disposable components disposed within one or more vessels (e.g., cart, holder) between the system 1200 and a predetermined location (e.g., across a room, in another room) separated from the system 1200. The vessel may be coupled to a drive mechanism configured to transport (e.g., self-propel) the vessel along the conveyor track. The conveyor tracks may be disposed along one or more surfaces of a room (e.g., wall, floor, ceiling). In some variations, each type of fluid may be transported to a corresponding location. For example, an unprocessed blood sample may be transported by the loader 1215 from a first room to a second room including the system 1200. Thereafter, an unprocessed (e.g., waste) blood sample may be transported by the loader 1215 from the system 1200 in the second room to a waste sample reservoir in a third room. Different types of fluids and / or materials may transported along different conveyor tracks and / or routes to corresponding fluid / material stations. In this manner, human interaction and / or contamination of a blood sample may be minimized and controlled, thereby enhancing sample tracking, automation, and high throughput processing.
[0306] In some variations, the manifold 1220 may be configured to couple to at least one microfluidic chip 1216. In some variations, the negative pressure source 1222 may be configured to couple to the manifold 1220. In this manner, a single manifold 1220 may be fluidically coupled to a plurality of microfluidic chips 1216 to facilitate application of negative pressure (e.g., vacuum, suction) to the plurality of microfluidic chips 1216 using a single negative pressure source 1222.
[0307] In some variations, the sensor 1224 may comprise one or more sensors configured to measure one or more characteristics (e.g., pressure, flow rate, optical image, temperature, humidity) corresponding to one or more of the biological fluid and components of the system 1200, 1300. In some variations, the input device 1226 may be configured to generate an input signal based on an operator input. In some variations, the processor 1228 and memory 1230 mayAttorney Docket No.: AUMI-004 / 02WO 348385-2089 be configured to control the system 1200, 1300. In some variations, the communication device 1232 may be configured to communicate with one or more components of the system 1200, 1300 as well as with networks and other computer systems. In some variations, the output device 1234 may be configured to output data corresponding to the system 1200, 1300 such as images of the biological fluid within the microfluidic chip 1216, flow rates, and the like.
[0308] In some variations, the system 1200, 1300 may be configured to perform multiple assays such as ELISA, immunoassay, microscopy, immunohistochemistry, fluorescence in situ hybridization, immunofluorescence, infrared, UV-VIS, Raman, NMR, mass spectrometry, IR spectroscopy, NG sequencing, protein array, ribonucleic acid array, gene array, qPCR, RT- qPCR, RT-PCR, and the like. In some variations, the microfluidic chip 1216 may be analyzed by the system 1200, 1300 and / or removed and analyzed using another device (e.g., multi-well plate, microscope slide). Examples of imaging techniques include electron microscopy, stereoscopic microscopy, wide-field microscopy, bright-field microscopy, phase-contrast microscopy, polarizing microscopy, phase contrast microscopy, multiphoton microscopy, differential interference contrast microscopy, fluorescence microscopy, laser scanning confocal microscopy, multiphoton excitation microscopy, ray microscopy, ultrasonic microscopy, positron emission tomography, computerized tomography, and magnetic resonance imaging.
[0309] FIGS.13A and 13B are perspective views of a system 1300 comprising a holder 1312, a reservoir 1313, a robot 1314, an end effector 1315 (e.g., pipette) coupled to the robot 1314, a plurality of chip connectors 1318 with each chip connector 1318 holding a plurality microfluidic chips 1316, a manifold 1320 including a plurality of fluid conduits 1321, and a negative pressure source 1322. The system 1300 may include a sensor, an input device, a processor, a memory, a communication device, and an output device are not shown in FIGS.13A-13B for the sake of clarity. In some variations, the end effector 1315 may comprise a plurality of pipettes configured to transfer fluid(s) stored in the holder 1312 to the microfluidic chips 1316 held in respective chip connectors 1318. In some variations, the system 1300 may include a plurality of robots 1314. For example, a first robot may be configured to transfer biological fluids, a second robot (not shown for the sake of clarity) may be configured to transfer reagents, and a third robot (not shown for the sake of clarity) may be configured to releasably couple the microfluidic chips 1316 from respective chip connectors 1318. The third robot may be configured to assemble and disassemble the disposable components including the microfluidic chips 1316 from the durableAttorney Docket No.: AUMI-004 / 02WO 348385-2089 components (e.g., chip connector 1318) of the system 1300 to reduce manual labor and / or reduce contamination. In FIG.13A, each microfluidic chip 1316 is coupled to a respective fluid conduit 1321, each of which is coupled to the manifold 1320. A single fluid conduit may couple the manifold 1320 to the reservoir 1313 and negative pressure source 1321.
[0310] In some variations, one or more fluids may be loaded from the holder 1312 to a respective inlet of a plurality of microfluidic chips 1316 using the robot 1314. The negative pressure source 1322 may apply suction to each of the microfluidic chips 1316 via the manifold 1320 and fluid conduits 1321. The fluid at the inlet (e.g., inlet reservoir) may be pulled toward the respective outlet of the microfluidic chip 1316 while ECMBs remain in the microfluidic chip 1316. The separated non-ECMB fluid is pulled through the fluid conduits 1321 and manifold 1320 and received at the reservoir 1313 (e.g., waste disposal).
[0311] The inlet connector 1340 may be coupled between an inlet 1330 of the microfluidic chip 1316 and one or more of the end effector 1315 and the funnel 1360. In some variations, the funnel 1360 may be configured to receive the end effector 1315 using the robot 1314. For example, the funnel 1360 may be configured to receive and / or guide fluid being transferred to the microfluidic chip 1316 from the end effector 1315. The outlet connector 1342 may be coupled between an outlet 1332 of the microfluidic chip 1316 and the fluid conduit 1321. In some variations, the fluid conduit 1321 may be configured to receive fluid being transferred from the microfluidic chip 1316 to the reservoir 1313 (e.g., via negative pressure).
[0312] In some variations, one or more clamps 1350 may be configured to releasably couple (e.g., secure, hold) the inlet connector 1340, respectively, and the outlet connector 1342 to the microfluidic chip 1316. For example, the clamp 1350 may comprise one or more springs (not shown) and a hinge 1352 configured to provide a predetermined force to hold the respective inlet connector 1340 and outlet connector 1342 in place relative to the microfluidic chip 1316. In some variations, a portion 1354 (e.g., actuator) of the clamp 1350 may be actuated (e.g., pushed down by a robot or operator) to release one or more of the inlet connector 1340 and the outlet connector 1342 to facilitate removal of the microfluidic chip 1316 from the chip connector 1318.Attorney Docket No.: AUMI-004 / 02WO 348385-2089 A. Holder
[0313] The systems described herein may comprise a holder 1212, 1312 configured to receive one or more fluids. In some variations, the holder 1212, 1312 may be configured to receive and store a plurality of fluids including a biological fluid (e.g., sample) and one or more reagents. Examples of reagents include a buffer, a lysing solution, a nucleic acid cleavage agent, a cleavage inhibitor, a precipitation agent, a fixative reagent, a carrier fluid, a biofluid, water, purified water, a saline solution, an organic solvent, a gelling agent, a surfactant, a ligand for binding or associating with a component of an ECMB, combinations thereof, and reagents for interacting with biological components of the sample. In some variations, the reagent may include one or more reagents for measuring a biomarker level or quantity, or for comparing a biomarker level to a control. In some variations, the holder 1212, 1312 may comprise a plurality of reservoirs to separately store each fluid without mixing. Fluids may include any suitable fluids, including for example biological fluids for sampling, and / or one or more fluids useful in analysis. Biological fluids for sampling may include, for example, one or more of human or animal bodily fluids, tissue, cells, whole blood, blood plasma, blood serum, cerebrospinal fluid, intrathecal fluid, urine, saliva, sweat, tears, synovial fluid, pleural fluid, gastric fluid, peritoneal fluid, breast milk, nipple aspirate, semen, amniotic fluid, vitreous, aqueous humor, lymph, bile, cerumen, chyle, chyme, endolymph, perilymph, exudates, feces, ejaculate, gastric acid, gastric juice, mucus, pericardial fluid, pus, rheum, sebum, serous fluid, smegma, sputum, synovial fluid, vaginal secretion, menstrual effluent, vomit, tumors, carriers, reagents, solutions, binding moieties, and the like. Biological fluids useful in analysis may include carriers, reagents, binding moieties, and the like. The holder 1212, 1312 may be disposed within the system 1200, 1300 at a location within reach of the robot 1214, 1314. For example, the holder 1212, 1312 may be configured to receive a plurality of pipettes coupled to the robot 1214, 1314 for transfer of a plurality of fluids from the holder 1212, 1312 to a plurality of microfluidic chips 1216, 1316. In some variations, a portion of the holder 1212 (e.g., vessel) containing a fluid (e.g., subject sample) may comprise a unique identifier (e.g., barcode, QR code, label) configured to facilitate tracking and identification of the fluid as it is processed through the system 1200. The unique identifiers described herein may be further associated with one or more of patient identification information, sample tracking information, clinical data, and the like. A sensor 1224 (e.g., barcode reader) may be configured to read the unique identifier at predetermined intervals and / or locations within the system 1200.Attorney Docket No.: AUMI-004 / 02WO 348385-2089 B. Robot
[0314] The systems described herein may comprise one or more robots 1214, 1314. Generally, a robot 1214, 1314 may be configured to move and transfer one or more fluids (e.g., biological fluid, reagent, antibodies, buffer) between a holder 1212, 1312 and at least one microfluidic chip 1216, 1316. For example, the robot 1214, 1314 may be configured to automatically load at least one microfluidic chip 1216, 1316 with fluids absent manual (e.g., human) intervention. In some variations, the robot 1214, 1314 may be coupled to an end effector comprising one or more pipettes (e.g., micro-pipette, multi-channel pipette) configured to transfer a fluid (e.g., biological fluid, reagent) from the holder 1212, 1312 to at least one microfluidic chip 1216, 1316 (e.g., through an inlet connector of a chip connector 1218, 1318). For example, a biological fluid from a subject may be first transferred from the holder 1212, 1312 to a plurality of microfluidic chips 1216, 1316 using the robot 1214, 1314. Then, different reagents may be transferred from the holder 1212, 1312 to predetermined microfluidic chips 1216, 1316 using the robot 1214, 1314 for different processing (e.g., different histochemical stains). The robot 1214, 1314 may be coupled to the processor 1228 and memory 1230 to control a type and volume of fluid transferred using one or more of the pipettes. In some variations, the end effector may be configured to releasably couple (e.g., assemble, disassemble) the disposable components from the durable components of the system. Additionally or alternatively, the system may comprise one or more material transport components such as a track and container configured to translate along the track.
[0315] In some variations, the robot 1214, 1314 may be configured to removably couple a microfluidic chip 1216, 1316 to one or more of a chip connector 1218, 1318 and a negative pressure source 1222, 1322. For example, the robot 1214, 1314 may be configured to couple a microfluidic chip 1216, 1316 to a chip connector 1218, 1318 such that the microfluidic chip 1216, 1316 is fluidically coupled to the manifold 1220, 1320 and negative pressure source 1222, 1322. Furthermore, robot 1214, 1314 may be configured to releasably couple other components (e.g., inlet connector, outlet connector) to the microfluidic chip 1216, 1316. Conversely, the robot 1214, 1314 may be configured to de-couple the microfluidic chip 1216, 1316 from the chip connector 1218, 1318 for transfer to another device (e.g., microscope slide, imaging system, analysis system). In some variations, the robot 1214, 1314 may be configured to remove one or more of the holder 1212, 1312 and reservoir 1213, 1313 from the system 1200, 1300 (e.g., toAttorney Docket No.: AUMI-004 / 02WO 348385-2089 replace with a different holder 1212, 1312 and reservoir 1213, 1313). Automated system set-up, fluid transfer, and cleaning using the robot 1214, 1314 may, among other benefits, reduce contamination and human error, and increase consistency and throughput.
[0316] In some variations, the robot 1214, 1314 may be, for example, a linear robot arm, an articulated robotic arm, and / or a SCARA robotic arm. The robot 1214, 1314 may comprise one or more segments coupled together by a joint (e.g., shoulder, elbow, wrist) configured to provide a single degree of freedom. Joints are mechanisms that provide a single translational or rotational degrees of freedom. For example, the robot 1214, 1314 may have six or more degrees of freedom. The set of Cartesian degrees of freedom may be represented by three translational (position) variables (e.g., surge, heave, sway) and by the three rotational (orientation) variables (e.g., roll, pitch, yaw). In some variations, the robot 1214, 1314 may have less than six degrees of freedom.
[0317] In some variations, the robot 1214, 1314 may be configured to move over all areas of a system 1200, 1300 in up to three dimensions. The robot 1214, 1314 may comprise one or more motors configured to translate and / or rotate the joints and move the robot 1214, 1314 to a desired location and orientation. In some variations, the position of the robot may be temporarily locked when delivering fluid to a predetermined microfluidic chip or when one or more microfluidic chips 1216, 1316 are being removed from a chip connector 1218, 1318 (e.g., for transfer to a microscope or other imaging system). The robot 1214, 1314 may be mounted to any suitable object, such as a platform (e.g., table), a wall, a ceiling, or may be self-standing (e.g., on the ground). Additionally or alternatively, the robot 1214, 1314 may be configured to be moved manually. C. Microfluidic chip
[0318] The systems described herein may comprise one or more microfluidic chips 1216, 1316. Generally, a microfluidic chip 1216, 1316 may be configured to receive and process one or more fluids. In some variations, a microfluidic chip 1216, 1316 may comprise an inlet reservoir, at least one channel (e.g., restriction channel, uniform flow channel), at least one obstruction (e.g., pillar) configured to restrict fluid flow, and an outlet reservoir. In some variations, one or more restriction channels may be configured to facilitate laminar flow within the channel. A restriction channel configured for laminar flow may improve consistency ofAttorney Docket No.: AUMI-004 / 02WO 348385-2089 sample processing. For example, processing a sample using laminar flow may ensure that samples associated with the same cancer parameter (e.g., CRC, PDAC, healthy control) produce similar morphological features and local spatialization patterns when processed on the microfluidic chip. Consistent processing of samples to produce repeatable features and / or patterns in separated ECMBs of the sample may improve the accuracy of the cancer parameter prediction methods described herein.
[0319] The restriction channel may include an inlet and an outlet. The inlet reservoir may be fluidically coupled to the inlet of the restriction channel, and the outlet of the restriction channel may be fluidically coupled to the outlet reservoir. The inlet reservoir may be configured to receive and store fluid transferred from robot 1214, 1314. The microfluidic chip 1216, 1316 may be configured to process a biological fluid through the channel while receiving negative pressure from an outlet (e.g., outlet reservoir) of the microfluidic chip 1216, 1316. Microfluidic chips 1216, 1316 having a plurality of channels enable higher throughput processing and reduced costs and system size. For example, a microfluidic chip 1216, 1316 having 8 channels may enable ECMB separation for 8 fluid samples for a plurality of histology stains, immunohistochemical (IHC) stains, and the like.
[0320] FIG.14 is a plan view of a microfluidic chip 1416 comprising an inlet reservoir 1410, an outlet reservoir 1430, and a restricted region 1420 (e.g., filter region) in fluid communication with the inlet reservoir 310 and the outlet reservoir 1430. The inlet reservoir 1410 may include an inlet 1412 (e.g., opening) and at least one obstruction 1440 (e.g., pillar). In some variations, the obstruction 1440 may have one or more of a circular, a spherical, a triangular, a square, a polygonal, a diamond, and a fin-shape. The restricted region 1420 may include a flow barrier 1422 and one or more restriction channels including a middle channel 1424, a top channel 1426, and bottom channel 1428. Likewise, the outlet reservoir 1430 may include an outlet 1432 (e.g., opening) and at least one obstruction 1440. The size, shape, and spacing of the obstructions 340 in the inlet reservoir 310 and outlet reservoir 330 may be the same or different. In some variations, fluid may be received at inlet 1412 and a fluid conduit coupled to a manifold and a negative pressure source may be coupled to the outlet 1432.
[0321] In some variations, the restricted region 1420 may be configured to restrict fluid flow so as to separate and hold (e.g., trap) a first fraction of the fluid within the restricted region 1420 of the microfluidic chip1416 while a remaining second fraction of the fluid is permitted to flowAttorney Docket No.: AUMI-004 / 02WO 348385-2089 out of the microfluidic chip 1416 via the negative pressure applied to the microfluidic chip 1416. In some variations, the first fraction of the fluid may include ECMBs. As shown in FIG.14, the restricted region 1420 may include flow barriers 1422 that define the restriction channel 1424 (e.g., uniform flow channel) and a plurality of obstructions 1440 (e.g., pillars) having a plurality of spacings 1426. The restriction channel 1424 may be linear and have a length at least equal to a length of the restricted region 1420. In some variations, one or more of laminar flow and flow rate consistency may be based on a length of the restriction channel 1424. For example, a restriction channel 1424 that does not extend into one or more of the inlet reservoir 1410 and the outlet reservoir 1430 may increase one or more of laminar flow and flow rate consistency throughout the microfluidic chip (e.g., along a width and length of the microfluidic chip 1416).
[0322] The plurality of obstructions 1440 within the restricted region 1420 may be configured to restrict (e.g., impede) fluid flow and hold the first fraction of the biological fluid. For example, a spacing between the plurality of obstructions 1440 in the restricted region 1420 may decrease along a length of the microfluidic chip 1416 from an inlet of the restricted region 1420 to an outlet of the restricted region1420. In some variations, the spacing between the plurality of obstructions 1440 in the restricted region 1420 may be between about 100 μm and about 4 μm. For example, the spacing between obstructions 1440 at an inlet of the restricted region 320 may be about 100 μm whereas the spacing between obstructions 1440 at an outlet of the restricted region 320 may be about 4 μm. In some variations, the spacing may vary in stepwise increments. For example, the spacing may decrease in order from about 100 μm to about 50 μm to about 25 μm to about 15 μm and to about 4 μm. Additionally or alternatively, the spacing between obstructions 1440 may vary continuously along a length of the restricted region 1420. In some variations, a larger proportion of ECMBs may be captured within the portions of the restricted region 1420 having smaller spacing.
[0323] The restriction channel 1424 may have a length of between about 5 mm and about 30 mm, between about 10 mm and about 30 mm, between about 15 mm and about 30 mm, between about 20 mm and about 30 mm, between about 5 mm and about 25 mm, between about 5 mm and about 20 mm, about 15 mm, about 20 mm, about 25 mm, about 30 mm, including all ranges and sub-values in-between. The restriction channel 1424 may have a cross-sectional dimension of between about 5 μm and about 30 μm, between about 10 μm and about 30 μm, between about 15 μm and about 30 μm, between about 20 μm and about 30 μm, between about 5 μm and aboutAttorney Docket No.: AUMI-004 / 02WO 348385-2089 25 μm, between about 5 μm and about 20 μm, about 15 μm, about 20 μm, about 25 μm, about 30 μm, including all ranges and sub-values in-between. In some variations, the restriction channel 1424 may comprise at least one obstruction. The plurality of obstructions 1440 may comprise a diameter of about 50 μm and about 1 mm.
[0324] In some variations, a microfluidic chip array may comprise a plurality of microfluidic chips 1416. For example, a microfluidic chip array may comprise a set of eight parallel microfluidic chips 1416. In this manner, a single microfluidic chip array may be used to separately (e.g., individually) process a plurality of fluid samples, thereby increasing throughput and efficiency, as well as reducing a size of the system. For example, a microfluidic chip array having eight microfluidic chips 1416 may enable independent ECMB separation for eight fluid samples using any combination of histology stains, immunohistochemical (IHC) stains, reagents, and the like. In some variations, a microfluidic chip array may increase one or more of laminar flow and flow rate consistency throughout the microfluidic chips of the plurality of microfluidic chips.
[0325] In some variations, ECMB capture amounts using the negative pressure systems as described herein may depend on one or more microfluidic channel dimensions. In some variations, a microfluidic chip 1216, 1316, 1416 applied with negative pressure may improve a yield (e.g., immobilization) of a predetermined fraction of a biological fluid based on a cross- sectional dimension of a channel. For example, each channel may comprise a cross-sectional (e.g., height) dimension of between about 5 μm and about 30 μm, between about 5 μm and about 10 μm, between about 5 μm and about 20 μm, between about 5 μm and about 15 μm, between about 10 μm and about 20 μm, between about 15 μm and about 30 μm, including all ranges and sub-values in-between.
[0326] In some variations, the microfluidic chip 1216, 1316, 1416 may comprise a unique identifier (e.g., barcode, QR code, label) configured to facilitate tracking and identification of the microfluidic chip 1216, 1316, 1416. In some variations, the system may comprise an optical sensor (e.g., barcode reader) configured to read the unique identifier (e.g., barcode, QR code, label).Attorney Docket No.: AUMI-004 / 02WO 348385-2089 D. Chip connector
[0327] The systems described herein may comprise a chip connector. Generally, a chip connector 1218, 1318 may be configured to hold one or more microfluidic chips 1216, 1316, 1416 and improve performance of the system 1200, 1300. Application of negative pressure to a microfluidic chip generates forces on the structure of the microfluidic chip itself that may reduce and / or prevent processing of a biological fluid using the microfluidic chip. For example, negative pressure applied to a channel (e.g., restriction channel) of a microfluidic chip may generate force at an outlet that may compress (and / or collapse) an outlet reservoir of the microfluidic chip corresponding to unpredictable flow rates, air bubbles, and unequal distribution of material within the microfluidic chip. By contrast, the chip connectors described herein may be configured to hold a microfluidic chip and distribute the negative pressure forces applied by a negative pressure source more evenly throughout a microfluidic chip 1216, 1316, 1416 so as to maintain the structural integrity of the microfluidic chip 1216, 1316, 1416 and provide consistent results.
[0328] In some variations, a geometry of the disposable inlet and outlet connectors described herein may improve operational aspects of a negative pressure system such as assembly and cleaning that may reduce a set-up time of the system. In some variations, an overall experiment duration for exemplary staining procedures may be improved using a negative pressure system relative to a positive pressure system. In some variations, the improved fluid seal of the disposable inlet and outlet connectors described herein may reduce contamination in a microfluidic chip corresponding to abnormal morphology relative to a reusable connector used in a positive pressure system. E. Manifold
[0329] The systems described herein may comprise a manifold 1220, 1320. Generally, a manifold 1220, 1320 may be configured to couple to a plurality of microfluidic chips 1216, 1316, 1416 and negative pressure source 1222, 1322. For example, the manifold 1220, 1320 may be in fluid communication with a single negative pressure source 1222, 1322 and each microfluidic chip 1216, 1316, 1416 held by a chip connector 1218, 1318 such that a negative pressure generated by the negative pressure source 1222, 1322 may be applied to each microfluidic chip 1216, 1316, 1416 coupled to the manifold 1220, 1320. In some variations, theAttorney Docket No.: AUMI-004 / 02WO 348385-2089 manifold 1220, 1320 may comprise a plurality of fluid conduits (e.g., vacuum tubes, fluid lines) configured to be releasably coupled to each of the microfluidic chips 1216, 1316, 1416 and the negative pressure source 1222, 1322. Accordingly, the manifold may reduce the size of the system 1200 since each microfluidic chip does not need its own negative pressure source. In some variations, the manifold and negative pressure source significantly increase throughput by scaling and increasing efficiency without a reduction in consistency or flowrate. A microfluidic systems described herein comprising a manifold and negative pressure source help to not only increase efficiency and throughput, but do so in a relatively uniform manner. F. Negative Pressure Source
[0330] The systems described herein may comprise a negative pressure source 1222, 1322. Generally, a negative pressure source 1222, 1322 may be configured to couple to a manifold 1220, 1320 and at least one microfluidic chip 1216, 1316, 1416. Application of negative pressure may help reduce contamination and increase processing throughput. In some variations, the negative pressure source may comprise a fluid pump. The processor 1228 and memory 1230 coupled to the negative pressure source 1222, 1322 may be configured to control the fluid flow rate and negative pressure. In some variations, the negative pressure source 1222, 1322 may be configured to apply a negative pressure of between about 10 mm HG and about 760 mm HG to each microfluidic chip 1216, 1316, 1416 of the system 1200, 1300. G. Sensor
[0331] The systems described herein may comprise one or more sensors 1224. Generally, a sensor 1224 may be configured to measure one or more characteristics (e.g., pressure, flow rate, optical image, temperature, humidity) corresponding to one or more of the biological fluid, microfluidic chip 1216, 1316, 1416. In this way, the system and the fluids may be monitored during use. In some variations, a sensor may be coupled to or integrated into any component of the system 1200 such as the inlet connector, chip connector, etc.
[0332] In some variations, the sensor may be an optical sensor coupled to any component of the system 1200, 1300 such as the chip connector 1218, 1318. The optical sensor may be configured to image one or more channels of the microfluidic chips 1216, 1316, 1416 for histochemical and morphological study. The optical sensor may be used to receive light signalsAttorney Docket No.: AUMI-004 / 02WO 348385-2089 (e.g., light beams) reflected by the fluid in the microfluidic chip 1216, 1316, 1416. The received light may be used to generate signal data that may be processed by the processor 1228 and memory 1230 to generate sample data. The optical sensor may further be configured to image one or more identifiers (e.g., label, barcode) and identifiers of the microfluidic chip 1216, 1316, 1416. In some variations, the optical sensor may include one or more of a lens, camera, and measurement optics. For example, the optical sensor may include a charged coupled device (CCD) or complementary metal-oxide semiconductor (CMOS) optical sensor and may be configured to generate an image signal that is transmitted to a output device 1234 (e.g., display). For example, the optical sensor may include a camera with an image sensor (e.g., a CMOS or CCD array with or without a color filter array and associated processing circuitry). Additionally or alternatively, the sensor in some variations may be an ultrasound sensor configured to generate an ultrasound signal used to determine a fluid flow rate.
[0333] In some variations, the system 1200, 1300 may include a radiation source configured to emit a light signal (e.g., illumination) directed at the microfluidic chip 1216, 1316, 1416 for visualization, imaging, and / or cleaning (e.g., UV light). In some variations, the radiation source may include one or more of a light emitting diode, laser, microscope, optical sensor, lens, and flash lamp. For example, the radiation source may generate light that may be carried by fiber optic cables or one or more LEDs may be configured to provide illumination. In another example, a fiberscope including a bundle of flexible optical fibers may be configured to receive and propagate light from an external light source.
[0334] In some variations, one or more sensors 1224 may be located separately from the system 1200. For example, a processed microfluidic chip 1216 may be transported using the loader 1215 to an optical sensor located separately from the system 1200. H. Input device
[0335] Generally, an input device 1226 of a system 1200, 1300 may serve as a communication interface between an operator and the system 1200, 1300. The input device 1226 may be configured to receive input data and output data to one or more of the robot 1214, sensor 1224, and output device 1234. For example, operator control of an input device 1226 (e.g., foot controller, joystick, keyboard, touch screen) may be processed by processor 1228 and memory 1230 for input device 1226 to output a control signal to one or more of robot 1214, 1314,Attorney Docket No.: AUMI-004 / 02WO 348385-2089 negative pressure source, and sensor 1224. As another example, images generated by sensor 1224 may be processed by processor 1228 and memory 1230, and displayed by the output device 1234 (e.g., display). Sensor data from one or more sensors 1224 may be output visually, audibly, and / or through haptic feedback by one or more output devices 1234.
[0336] Some variations of an input device may comprise at least one switch configured to generate a control signal. The input device may be coupled to or separate from other components of the system 1200, 1300. For example, the input device 1226 may be disposed in a different room than the robot 1214, 1314 and microfluidic chip 1216, 1316, 1416 to reduce potential contamination. The control signal may include, for example, a robot signal, a negative pressure signal, a sensor signal, and other signals. In some variations, the input device 1226 may comprise a wired and / or wireless transmitter configured to transmit a control signal to a wired and / or wireless receiver of a controller. A robot signal (e.g., for the control of movement, position, and orientation) may control articulation of the robot in at least four degrees of freedom of motion, and may include yaw and / or pitch rotation. For example, an input device 1226 comprising a touch surface may be configured to detect contact and movement on the touch surface using any of a plurality of touch sensitivity technologies including capacitive, resistive, infrared, optical imaging, dispersive signal, acoustic pulse recognition, and surface acoustic wave technologies.
[0337] In variations of an input device 1226 comprising at least one switch, a switch may comprise, for example, at least one of a button (e.g., hard key, soft key), touch surface, keyboard, analog stick (e.g., joystick), directional pad, mouse, trackball, jog dial, step switch, rocker switch, pointer device (e.g., stylus), motion sensor, image sensor, and microphone. A motion sensor may receive operator movement data from an optical sensor and classify an operator gesture as a control signal. A microphone may receive audio and recognize an operator voice as a control signal. In variations of a system comprising a plurality of input devices, different input devices may generate different types of signals. For example, some input devices (e.g., button, analog stick, directional pad, and keyboard) may be configured to generate a robot signal while other input devices (e.g., step switch, rocker switch) may be configured to control a negative pressure source 1222, 1322 and the sensors 1224.Attorney Docket No.: AUMI-004 / 02WO 348385-2089 I. Processor
[0338] A system 1200, 1300, as depicted in FIG.12, may comprise a processor 1228 and a machine-readable memory 1230 (e.g., controller) in communication with the robot 1214, 1314, negative pressure source 1222, 1322, and sensor(s) 1224. The processor 1228 may be connected to the system 1200, 1300 by wired or wireless communication channels. The processor 1228 may be located in the same or different room as the microfluidic chips 1216, 1316, 1416. The processor 1228 may be configured to control one or more components of the system 1200, 1300, such as the robot 1214, 1314 configured to transfer fluids to the microfluidic chip or an optical sensor configured to visualize ECMBs separated using the microfluidic chip 1216, 1316, 1416.
[0339] The processor 1228 may be implemented consistent with numerous general purpose or special purpose computing systems or configurations. Various exemplary computing systems, environments, and / or configurations that may be suitable for use with the systems and devices disclosed herein may include, but are not limited to software or other components within or embodied on personal computing devices, network appliances, servers or server computing devices such as routing / connectivity components, portable (e.g., hand-held) or laptop devices, multiprocessor systems, microprocessor-based systems, and distributed computing networks.
[0340] Examples of portable computing devices include smartphones, personal digital assistants (PDAs), cell phones, tablet PCs, phablets (personal computing devices that are larger than a smartphone, but smaller than a tablet), wearable computers taking the form of smartwatches, portable music devices, and the like, and portable or wearable augmented reality devices that interface with an operator’s environment through sensors and may use head- mounted displays for visualization, eye gaze tracking, and user input.
[0341] The processor 1228 may incorporate data received from memory 1230 and operator input to control one or more robots 1214, 1314 and negative pressure source 1222, 1322. The memory 1230 may further store instructions to cause the processor 1228 to execute modules, processes, and / or functions associated with the system 1200, 1300. The processor 1228 may be any suitable processing device configured to run and / or execute a set of instructions or code and may comprise one or more data processors, image processors, graphics processing units, physics processing units, digital signal processors, and / or central processing units. The processor 1228 may be, for example, a general purpose processor, a Field Programmable Gate Array (FPGA), anAttorney Docket No.: AUMI-004 / 02WO 348385-2089 Application Specific Integrated Circuit (ASIC), configured to execute application processes and / or other modules, processes, and / or functions associated with the system and / or a network associated therewith. The underlying device technologies may be provided in a variety of component types such as metal-oxide semiconductor field-effect transistor (MOSFET) technologies like complementary metal-oxide semiconductor (CMOS), bipolar technologies like emitter-coupled logic (ECL), polymer technologies (e.g., silicon-conjugated polymer and metal- conjugated polymer-metal structures), mixed analog and digital, combinations thereof, and the like. J. Memory
[0342] Some variations of memory 1230 described herein relate to a computer storage product with a non-transitory computer-readable medium (also may be referred to as a non-transitory processor-readable medium) having instructions or computer code thereon for performing various computer-implemented operations. The computer-readable medium (or processor- readable medium) is non-transitory in the sense that it does not include transitory propagating signals per se (e.g., a propagating electromagnetic wave carrying information on a transmission medium such as air or a cable). The media and computer code (also may be referred to as code or algorithm) may be those designed and constructed for a specific purpose or purposes. Examples of non-transitory computer-readable media include, but are not limited to, magnetic storage media such as hard disks, floppy disks, and magnetic tape; optical storage media such as Compact Disc / Digital Video Discs (CD / DVDs), Compact Disc-Read Only Memories (CD- ROMs), and holographic devices; magneto-optical storage media such as optical discs; solid state storage devices such as a solid state drive (SSD) and a solid state hybrid drive (SSHD); carrier wave signal processing modules; and hardware devices that are specially configured to store and execute program code such as Application-Specific Integrated Circuits (ASICs), Programmable Logic Devices (PLDs), Read-Only Memory (ROM), and Random-Access Memory (RAM) devices. Other variations described herein relate to a computer program product, which may include, for example, the instructions and / or computer code disclosed herein.
[0343] The systems, devices, and / or methods described herein may be performed by software (executed on hardware), hardware, or a combination thereof. Software modules (executed on hardware) may be expressed in a variety of software languages (e.g., computer code), includingAttorney Docket No.: AUMI-004 / 02WO 348385-2089 C, C++, Java®, Python, Ruby, Visual Basic®, and / or other object-oriented, procedural, or other programming language and development tools. Examples of computer code include, but are not limited to, micro-code or micro-instructions, machine instructions, such as produced by a compiler, code used to produce a web service, and files containing higher-level instructions that are executed by a computer using an interpreter. Additional examples of computer code include, but are not limited to, control signals, encrypted code, and compressed code. K. Communication device
[0344] In some variations, systems 1200, 1300 described herein may communicate with networks and computer systems through a communication device 1228. In some variations, the system 1200, 1300 may be in communication with other devices via one or more wired and / or wireless networks. A wireless network may refer to any type of digital network that is not connected by cables of any kind. Examples of wireless communication in a wireless network include, but are not limited to cellular, radio, satellite, and microwave communication. However, a wireless network may connect to a wired network in order to interface with the Internet, other carrier voice and data networks, business networks, and personal networks. A wired network is typically carried over copper twisted pair, coaxial cable and / or fiber optic cables. There are many different types of wired networks including wide area networks (WAN), metropolitan area networks (MAN), local area networks (LAN), Internet area networks (IAN), campus area networks (CAN), global area networks (GAN), like the Internet, and virtual private networks (VPN). Hereinafter, network refers to any combination of wireless, wired, public and private data networks that are typically interconnected through the Internet, to provide a unified networking and information access system.
[0345] Cellular communication may encompass technologies such as GSM, PCS, CDMA or GPRS, W-CDMA, EDGE or CDMA2000, LTE, WiMAX, and 5G networking standards. Some wireless network deployments combine networks from multiple cellular networks or use a mix of cellular, Wi-Fi, and satellite communication. In some variations, the communication device 1232 may comprise a radiofrequency receiver, transmitter, and / or optical (e.g., infrared) receiver and transmitter. The communication device 1232 may communicate by wires and / or wirelessly with one or more of the robot 1214, 1314, negative pressure source 1222, 1322, sensor 1224, input device 1226, output device 1234, network, database, server, combinations thereof, and the like.Attorney Docket No.: AUMI-004 / 02WO 348385-2089 L. Output device
[0346] An output device 1234 of a system 1200, 1300 may be configured to output data corresponding to a system, and may comprise one or more of a display device, audio device, and haptic device. For example, a display device may allow an operator to view images of one or more microfluidic chips 1216, 1316, 1416 and the robot 1214, 1314. In some variations, an output device may comprise a display device including at least one of a light emitting diode (LED), liquid crystal display (LCD), electroluminescent display (ELD), plasma display panel (PDP), thin film transistor (TFT), organic light emitting diodes (OLED), electronic paper / e-ink display, laser display, and / or holographic display.
[0347] An audio device may audibly output fluid data, sensor data, system data, alarms and / or warnings. For example, the audio device may output an audible warning when sensor data (e.g., pressure, flow rate) falls outside a predetermined range or when a malfunction in a robot is detected. As another example, audio may be output when operator input is overridden by the system to prevent potential harm to the operator and / or system (e.g., robot collision, excessive negative pressure). In some variations, an audio device may comprise at least one of a speaker, piezoelectric audio device, magnetostrictive speaker, and / or digital speaker. In some variations, an operator may communicate to other users using the audio device and a communication channel. For example, the operator may form an audio communication channel (e.g., VoIP call) with a remote operator and / or observer.
[0348] A haptic device may be incorporated into one or more of the input and output devices to provide additional sensory output (e.g., force feedback) to the operator. For example, a haptic device may generate a tactile response (e.g., vibration) to confirm operator input to an input device (e.g., touch surface). Additionally or alternatively, haptic feedback may notify that an operator input is overridden by the system to prevent potential harm to the operator and / or system (e.g., robot collision, excessive negative pressure). III. Examples
[0349] For each of the examples described herein, the analyzed ECMB data corresponds to ECMBs separated from a biological fluid (e.g., human plasma) by a microfluidic chip. The separated ECMBs provide an increased signal-to-noise ratio in predicting a cancer parameterAttorney Docket No.: AUMI-004 / 02WO 348385-2089 (e.g., cancer, non-cancer, cancer type, cancer stage, pre-cancer). In some variations, the separated ECMBs on the microfluidic chip were characterized with LC-MS / MS based proteomics. For example, two early-stage and one late- stage prostate, lung, breast, and colorectal cancer plasma samples were compared against six non-cancerous control samples. As described in detail herein, gel electrophoresis and mass spectral analysis of colorectal cancer plasma ECMBs show distinct differences in protein expression for early- stage and late-stage cancer compared to a non-cancerous control. Moreover, gene ontology analysis of the proteomics-derived biochemical profile of microfluidic chip derived ECMBs indicates increased activity of biological pathways involved in negative regulation of endopeptidase activity. Gene Ontology (GO) analysis further indicates that whole pancreatic cancer plasma and corresponding isolates have an increased composition of cellular components related to extracellular matrix material and extracellular vesicles. Biomarkers
[0350] Exemplary protein biomarkers for pancreatic cancer are given in Table 1. Table 1: Protein biomarkers for pancreatic cancerAttorney Docket No.: AUMI-004 / 02WO 348385-2089Attorney Docket No.: AUMI-004 / 02WO 348385-2089Attorney Docket No.: AUMI-004 / 02WO 348385-2089Attorney Docket No.: AUMI-004 / 02WO 348385-2089
[0351] Moreover, the data in Table 2 shows that the isolated extracellular matrix bodies surprisingly contained disease-associated biomarkers such as SAA1, CRP, and APOL1, among others, in early-stage pancreatic ductal adenocarcinoma (PDAC) plasma. Table 2: Proteomic profiling of extracellular matrix bodies in early-stage PDAC plasmaAttorney Docket No.: AUMI-004 / 02WO 348385-2089Attorney Docket No.: AUMI-004 / 02WO 348385-2089Attorney Docket No.: AUMI-004 / 02WO 348385-2089
[0352] Moreover, the data in Table 3 shows that the isolated extracellular matrix bodies contained certain disease-associated biomarkers such as THBS1, SAAn, and KLKB1, among others, which were observed for the first time in ECMBs from late-stage PDAC plasma. Table 3: Proteomic profiling of extracellular matrix bodies in late-stage PDAC plasmaAttorney Docket No.: AUMI-004 / 02WO 348385-2089Attorney Docket No.: AUMI-004 / 02WO 348385-2089Attorney Docket No.: AUMI-004 / 02WO 348385-2089Attorney Docket No.: AUMI-004 / 02WO 348385-2089
[0353] Moreover, the data in Table 4 shows that the isolated extracellular matrix bodies contained certain disease-associated biomarkers such as CRP, SAAn, and CD5L, among others, which were observed for the first time in ECMBs. Table 4: Proteomic profiling of extracellular matrix bodies in PDAC plasmaAttorney Docket No.: AUMI-004 / 02WO 348385-2089Attorney Docket No.: AUMI-004 / 02WO 348385-2089Attorney Docket No.: AUMI-004 / 02WO 348385-2089
[0354] A comparison between a control isolate fraction and colorectal cancer isolate fractions are described below with respect to FIGS.15A-15C. FIG.15A is a mass density plot 1500 of a gel plot for a control isolate fraction 1520 and a stage II colorectal cancer isolate fraction 1522. FIG.15B is a mass density plot 1502 of a gel plot for a control isolate fraction 1520 and a stage III colorectal cancer isolate fraction 1524. In particular, the plots 1500 and 1502 show differences in protein profile between the control isolate fraction and respective colorectal cancer isolate fractions. A larger broad peak around the 80 pixel region in the cancer isolate fractions 1522, 1524 obscures the smaller distinct peaks (e.g., arrowhead) present in the control isolate fraction 1520. A sharp peak in the control isolate fraction 1520 at around 75 pixels is absent in both cancer isolate fractions 1522, 1524. These peaks correspond to proteins that differ in their expression in colorectal cancer and may correspond to biomarkers for cancer diagnosis. The data shows that isolating ECMBs increases the signal to noise ratio for finding disease biomarkers and targets associated with ECMB pathology.
[0355] FIG.15C are images of respective gel plots for a control isolate fraction 1504, a stage II colorectal cancer isolate fraction 1506, and a stage III colorectal cancer isolate fraction 1508. Sodium dodecyl sulfate–polyacrylamide gel electrophoresis (SDS-PAGE) from control isolate fraction 1504, stage II colorectal cancer isolate fraction 1506, and stage III colorectal cancer isolate fraction 1508 are shown aligned in FIG.15C to highlight a difference in band profile around a 250 kDa region. The control isolate fraction 1504 shows two distinct bands 1505 that are absent in the cancer isolate fractions 1507. The bands 1505 correspond to proteins that differ in protein expression relative to the colorectal cancer isolate fractions 1506, 1508 and may correspond to biomarkers for early cancer diagnosis. The data shows that isolating ECMBsAttorney Docket No.: AUMI-004 / 02WO 348385-2089 increases the signal to noise ratio for finding disease biomarkers and targets associated with ECMB pathology. Pre-Cancer
[0356] Conventional cancer liquid biopsy methods for detecting colorectal cancer (e.g., Shield) have an unacceptably low sensitivity for detecting pre-cancerous polyps in the colon. Conventional cancer liquid biopsy methods for detecting advanced adenoma (AA) have a sensitivity below about 20%. The poor performance of these conventional techniques is likely due to the low abundance and low signal-to-noise ratio of the test analyte. For example, circulating tumor DNA (ctDNA) or other genomic methods for detecting AA are very low in abundance and likely excreted into the blood circulation. By contrast, pre-cancer such as pre- cancerous colon and rectal polyps and advanced adenoma may be predicted using one or more of ECMB morphology, histological and immunofluorescence staining, and prediction machine learning models with high accuracy, alone and / or in combination thereof. For example, ECMB morphology may be used to detect advanced adenoma (AA), a pre-cancer polyp in the colon or rectum with characteristics that indicate a higher risk for progressing to cancer, in human plasma and compare to non-cancerous ECMBs.
[0357] AA ECMBs have a distinct morphology relative to non-cancerous ECMB samples and may have similar morphology as stage 1-4 CRC EMCBs. With respect to prediction machine learning models, transfer learning in the CRC data set may help the prediction machine learning model recognize the AA ECMB samples. Accordingly, prediction machine learning models were trained using AA ECMB and CRC ECMB images in one cohort and control images in another. For example, AA ECMBs, CRC ECMBs, and healthy ECMBs were isolated and processed on a microfluidic ship as described herein, stained with eosin, and imaged using methods described herein. The data set produced from the images comprised a control group (i.e., non-cancerous) of 308 samples and a disease group of 324 samples comprising both AA ECMBs samples, 82 samples, and CRC samples, 242 samples. An eosin-based prediction machine leaning model was trained to discriminate between the disease group and the control group. The accuracy of the eosin-based AA prediction model in discriminating between the AA ECMB samples and the control group samples was evaluated by measuring the sensitivity and specificity of the model in predicting AA. The samples making up the training and validationAttorney Docket No.: AUMI-004 / 02WO 348385-2089 data had no overlap with the samples used to test and validate the model and all samples were from different individual subjects.
[0358] Distinct ECMB morphology is present in pre-cancerous lesions and may allow for very early disease detection and characterization when processed using histological stains on a microfluidic chip. For example the binary classification plot 8800 of FIG.88, shows the eosin- based AA prediction machine learning model trained on AA ECMB morphology on the microfluidic chip and staining with histological dyes (eosin) described above may classify AA from healthy controls with high sensitivity. Plot 8800 shows that the eosin-based AA model, at an average (e.g., unoptimized) threshold, classified AA from control samples with a sensitivity of 100% and a specificity of 50.81%. Plot 8800 shows the eosin-based AA model predicted AA with a positive predictive value (PPV) is 35% and a negative predictive value (NPV) of 100%.
[0359] Furthermore, the sensitivity and specificity of the prediction machine learning model may be optimized based on the threshold that provides a predetermined test accuracy. In some variations, the threshold of a prediction machine learning model may a predetermined value. In other variations, the threshold may be a predetermined range. The threshold may be optimized to reduce the difference between the sensitivity and specificity performance of the model, thereby significantly improving the balance between the sensitivity and specificity of the model. For example, at the average threshold, the eosin-based AA prediction model may favor sensitivity over specificity, leading to misclassification of control samples as AA. To improve the specificity of the model, the threshold may be adjusted. Table 8900 of FIG.89 shows the performance of the exemplary eosin-based AA prediction model at various thresholds. For example, a threshold of 0.85 minimized the difference between the sensitivity and specificity of the model and was determined to be an optimal threshold. At a threshold of 0.85, the model performance showed a sensitivity of 82.35% and a specificity of 75.81%, thereby demonstrating higher accuracy and sensitivity compared to conventional technologies. A range of thresholds from 0.15 to 0.95 demonstrated acceptable sensitivity and specificity.
[0360] Similarly, distinct ECMB morphology stained with immunofluorescence enriched antibodies (e.g., DAPI stain) may allow for very early disease detection and characterization. For example, AA ECMBs, CRC ECMBs, and healthy ECMBs were isolated and processed on a microfluidic ship as described herein. The samples containing the ECMBs were stained with talin antibody staining, Rab27b antibody staining, fibronectin antibody staining, zyxin antibodyAttorney Docket No.: AUMI-004 / 02WO 348385-2089 staining, and DAPI fluorescence staining. The ECMBs were imaged on the microfluidic chip using methods described herein. The data set used to train the IF-based AA prediction model comprised merged images produced from the images of each IF staining and brightfield imaging. The data set comprised a control group (i.e., non-cancerous) of 239 samples and a disease group of 284 samples comprising both AA ECMBs samples, 65 samples, and CRC samples, 219 samples. An IF-based prediction machine learning model was trained to discriminate between the disease group and the control group and the accuracy of the IF-based AA prediction model in discriminating between the AA ECMB samples and the control group samples was evaluated by measuring the sensitivity and specificity of the model in predicting AA. The samples making up the training and validation data had no overlap with the samples used to test and validate the model and all samples were from different individual subjects.
[0361] Distinct ECMB morphology is present in pre-cancerous lesions and may allow for very early disease detection and characterization using immunofluorescence antibody staining on a microfluidic chip. For example the binary classification plot 9000 of FIG.90, shows the prediction machine learning model trained on AA ECMB morphology on the microfluidic chip and staining with IF stains as described above may classify AA from healthy controls with high sensitivity. Plot 9000 shows that the model, at an average threshold (e.g., unoptimized threshold), classified AA from control samples with a sensitivity of 91.18% and a specificity of 54.55%. Plot 9000 shows the model predicted AA with a positive predictive value (PPV) is 40.79% and a negative predictive value (NPV) of 94.74%. Furthermore, the sensitivity and specificity of the IF-based AA prediction machine learning model was optimized to reduce the difference between the sensitivity and specificity to improve performance of the model. At the average threshold, the IF-based AA prediction model favored sensitivity over specificity, leading to misclassification of control samples as AA. To improve the specificity the threshold may be adjusted. Table 9100 of FIG.91 shows the performance of the exemplary IF-based AA prediction model at various thresholds. A threshold of 0.80 minimized the difference between the sensitivity and specificity of the IF-based AA model and was determined to be an optimal threshold. At a threshold of 0.80, the model performance showed a sensitivity of 76.47% and a specificity of 73.74% thereby demonstrating higher accuracy and sensitivity compared to conventional technologies. A range of thresholds from 0.15 to 0.90 demonstrated acceptable sensitivity and specificity for the IF-based AA prediction model.Attorney Docket No.: AUMI-004 / 02WO 348385-2089 Prediction machine learning model
[0362] The prediction machine learning models described herein may be configured to analyze ECMB morphology to predict a cancer parameter (e.g., classify a sample between non-cancerous and cancer) with high accuracy and precision. Non-cancerous control plasma ECMBs on a microfluidic chip demonstrate low variability in morphological features. However, cancerous plasma ECMBs (e.g., prostate cancer, colorectal cancer, pancreatic cancer, lung cancer, breast cancer) on a microfluidic chip may include morphologically distinct features that allow for cancer parameter prediction with high sensitivity, specificity, PPV, and NPV. For example, the prediction machine learning model may comprise one or more neural networks and data augmentation, enhanced by transfer learning, regularization techniques, and dropout to prevent overfitting and improve generalization. The prediction machine learning model demonstrated an average precision of 97.3% in differentiating between prostate cancer and control images, an average precision of 94.3% for classifying colorectal versus lung cancer, and 95.0% accuracy in identifying colorectal cancer from non-cancerous controls. In some variations, the prediction machine learning model may differentiate between cancer images and control images with an average precision of above about 80%, above about 85%, above about 90%, and above about 95%. In some variations, the prediction machine learning model may classify between cancer types with an accuracy of above about 80%, above about 85%, above about 90%, and above about 95%. In some variations, the prediction machine learning model may identify cancer from a non-cancerous control with an accuracy of above about 80%, above about 85%, above about 90%, and above about 95%.
[0363] FIG.16A is a binary classification plot 1600 for cancer prediction using Eosin staining corresponding to the prediction machine learning models configured to predict a cancer parameter including cancer type. For example, plot 1600 includes sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV) between prostate cancer, lung cancer, colorectal cancer, and a respective control sample. The prediction model classifying between prostate cancer and non-cancerous (e.g., control) demonstrated a sensitivity of 100%, a specificity of 80%, a PPV of 83.3%, and an NPV of 100% based on 111 total samples including 89 training samples, 12 validation samples, and 10 testing samples. The prediction model classifying between lung cancer and colorectal cancer, where the lung class is a positive and colorectal class is a negative, demonstrated a sensitivity of 75%, a specificity of 100%, a PPV ofAttorney Docket No.: AUMI-004 / 02WO 348385-2089 100%, and an NPV of 83.3% based on 96 total samples including 77 training samples, 10 validation samples, and 9 testing samples. The prediction model classifying between colorectal cancer and health (e.g., control) demonstrated a sensitivity of 100%, a specificity of 100%, a PPV of 100%, and an NPV of 100% based on 114 total samples including 86 training samples, 14 validation samples, and 14 testing samples.
[0364] FIG.16B is a confidence score table 1610 (e.g., true positive prediction confidence score) corresponding to a prediction machine learning model configured to predict a cancer parameter. In particular, table 1610 show the confidence scores for true positive predictions which range from moderate to high across the various classifications. For example, the prediction score has a range between 0.545 and 0.984 for prostate cancer. The prediction score has a range between 0.433 and 0.909 for a control (e.g., non-cancerous tissue). In classifying between colorectal cancer and lung cancer, the prediction machine learning model demonstrates a high degree of confidence with the prediction score having a range between 0.538 and 0.986 for colorectal cancer and between 0.421 and 0.998 for lung cancer. In classifying between colorectal cancer and a control, the prediction machine learning model demonstrates confidence in predicting colorectal cancer with the prediction score having a range between 0.628 and 0.735, and reasonable confidence in predicting control cases with the prediction scores having a range between 0.519 and 0.821. In some variations, the prediction machine learning model may predict cancer with a prediction score above about 0.5, above about 0.6, above about 0.7, above about 0.8, above about 0.9. In some variations, the prediction machine learning model may predict control cases with a prediction score above about 0.5, above about 0.6, above about 0.7, above about 0.8, above about 0.9.
[0365] The prediction machine learning models described herein may predict between non- cancerous tissue (e.g., control), lung cancer, pancreatic cancer, colorectal cancer, and prostate cancer. For example, FIG.17A is a binary classification plot for cancer prediction using Eosin staining corresponding to the prediction machine learning models described herein. In particular, the prediction model demonstrated a sensitivity of 83.3% and a PPV of 90% for lung cancer, a sensitivity of 100% and a PPV of 83.3% for pancreatic cancer, a sensitivity of 85.7% and a PPV of 100% for a control sample, a sensitivity of 83.3% and a PPV of 71.4% for colorectal cancer, and a sensitivity of 83.3% and a PPV of 100% for prostrate cancer. based on 251 total samples including 188 training samples, 32 validation samples, and 31 testing samples.Attorney Docket No.: AUMI-004 / 02WO 348385-2089
[0366] FIG.17B is a confusion matrix (e.g., error matrix) 1710 for cancer type classification using eosin staining corresponding to the prediction machine learning models described herein. Generally, a confusion matrix compares the number of actual class instances to the number of predicted class instances. In particular, the rows correspond to the true labels (e.g., actual conditions), while the columns correspond to the predicted labels (e.g., model predictions). The gray cells indicate the percentage of correct predictions for each true label, which are the diagonal cells from the top-left to the bottom-right. The other cells in the matrix represent the percentage of misclassifications. The prediction model demonstrated 100% accuracy in classifying pancreatic cancer plasma ECMBs and distinguishing it from control plasma ECMBs, lung cancer plasma ECMBs, prostate cancer plasma ECMBs, and colorectal cancer plasma ECMBs. Similarly, the prediction model demonstrated 100% accuracy in classifying colorectal cancer plasma ECMBs.
[0367] As described in more detail herein (e.g., and with respect to FIGS.29, 31, 32, 33A, 33B, 34, 35A, 35B, and 36), cancer prediction using the prediction machine learning models described herein provided unexpected results that are significantly improved relative to conventional methods of predicting cancer (e.g., ctDNA). For example, FIG.29 is a binary classification plot 2900 for colorectal cancer prediction using a prediction machine learning model as described herein based on 198 total samples (99 colorectal, 99 non-cancerous). The machine learning model received preprocessed images and was trained on using a variety of image configurations (e.g., a combined image including the top and bottom channels). The colorectal samples studied included participants with various stages of colorectal cancer including 23.2% having stage 1 colorectal cancer, 40.4% havin...
Claims
Attorney Docket No.: AUMI-004 / 02WO 348385-2089 CLAIMS We Claim:
1. A method of predicting cancer, comprising: receiving image data corresponding to extracellular matrix bodies (ECMBs) separated from a biological fluid of a subject; generating ECMB data based on the image data; and predicting a cancer parameter based on at least the ECMB data.
2. The method as in any preceding claim, further comprising confirming the predicted cancer parameter.
3. The method as in any preceding claim, wherein predicting the cancer parameter comprises analyzing the ECMB data using a prediction machine learning model.
4. The method as in any preceding claim, further comprising confirming the predicted cancer parameter based on the ECMB data analyzed using a prediction machine learning model.
5. The method as in any preceding claim, wherein predicting the cancer parameter comprises a first prediction by one of a human and a prediction machine learning model and a second prediction using an other of the human and the prediction machine learning model.
6. The method as in any preceding claim, wherein the ECMB data comprises one or more morphological features.
7. The method as in any preceding claim, wherein the one or more morphological features correspond to one or more of a shape, a size, a spatial location, an intensity, an abundance, a color, a texture, an edge property and a stiffness of the ECMB data.
8. The method as in any preceding claim, further comprising identifying a region of interest in the received image data comprising the one or more morphological features.Attorney Docket No.: AUMI-004 / 02WO 348385-2089 9. The method as in any preceding claim, wherein the region of interest is manually identified.
10. The method as in any preceding claim, wherein generating the ECMB data comprises generating a plurality of image parameters from the region of interest, wherein the image parameters correspond to one or more of image intensity, image color, image texture, image edge features, image wavelet features, and image statistical parameters.
11. The method as in any preceding claim, wherein predicting the cancer parameter comprises analyzing the plurality of image parameters using a prediction machine learning model.
12. The method as in any preceding claim, further comprising selecting one more image parameters of the plurality of image parameters based on one or more of the morphological features using the prediction machine learning model.
13. The method as in any preceding claim, wherein confirming the predicted cancer parameter is based on the plurality of image parameters.
14. The method as in any preceding claim, further comprising training a prediction machine learning model to analyze the ECMB data on unlabeled image parameter data.
15. The method as in any preceding claim, further comprising processing the ECMBs using a first stain and a second stain different from the first stain, wherein predicting the cancer parameter is based on the ECMBs corresponding to the first stain and confirming the predicted cancer parameter is based on the ECMBs corresponding to the second stain.
16. The method as in any preceding claim, wherein the first stain and the second stain comprise one or more of a protease inhibitor, a histochemical stain, a hematoxylin stain, an eosin stain, a histology stain, an alcian blue stain, a picrosirius red stain, an immunohistochemical (IHC) stain, immunofluorescent (IF) stain, a multiplex IHC or IF stain, multi-spectral imaging, a protein stain, glycomic stain, a nucleic acid stain, and chemical fixation.Attorney Docket No.: AUMI-004 / 02WO 348385-2089 17. The method as in any preceding claim, wherein generating the ECMB data comprises segmenting the received image data.
18. The method as in any preceding claim, wherein the ECMB data is segmented using the prediction machine learning model.
19. The method as in any preceding claim, further comprising training a prediction machine learning model to segment the received image data based on a manually identified region of interest.
20. The method as in any preceding claim, further comprising training a prediction machine learning model on one or more of image data corresponding to ECMBs, ECMB data, clinical data, and an electronic medical record.
21. The method as in any preceding claim, wherein the one or more morphological features are segmented.
22. The method as in any preceding claim, further comprising generating a confidence value of the predicted cancer parameter.
23. The method as in any preceding claim, wherein predicting the cancer parameter comprises a first prediction using a sensitivity-based prediction machine learning model and a second prediction using a specificity-based machine learning model.
24. The method as in any preceding claim, wherein predicting the cancer parameter comprises a first prediction using a specificity-based prediction machine learning model and a second prediction using a sensitivity-based machine learning model.
25. The method as in any preceding claim, wherein analyzing the ECMB data using a prediction machine learning model comprises a first prediction machine learning modelAttorney Docket No.: AUMI-004 / 02WO 348385-2089 associated with a first stain or a first biomarker and a second prediction machine learning model associated with a second stain or a second biomarker.
26. The method as in any preceding claim, wherein the cancer parameter is predicted as non- cancerous when the confidence value is below a predetermined threshold.
27. The method as in any preceding claim, wherein the prediction machine learning model comprises one or more of ResNet, Inception, boosting algorithms, bootstrap aggregation, random forests, decision trees, Vision Transformers, convolutional neural networks (CNNs), and combinations thereof.
28. The method as in any preceding claim, wherein predicting the cancer parameter comprises analyzing the ECMB data using at least one scoring rubric defining at least one feature set comprising one or more morphological features.
29. The method as in any preceding claim, wherein the at least one scoring rubric defines a first feature set associated with cancer and assigns the first feature set associated with cancer a first point value, and the at least one scoring rubric defines a second feature set associated with a non-cancerous sample and assigns the second feature set associated with the non-cancerous sample a second point value different from the first point value.
30. The method as in any preceding claim, further comprising excluding an erroneous sample defined by a third feature set of the at least one scoring rubric.
31. The method as in any preceding claim, wherein the cancer parameter is predicted based on a total point value.
32. The method as in any preceding claim, further comprising processing the biological fluid using one or more of microfluidic separation, chemical fixation, physical fixation, affinity chromatography, sedimentation, centrifugation, differential centrifugation, density gradient centrifugation, mesh filtration, diafiltration, tangential flow filtration, membrane filtration, elutriation, affinity-based capture, precipitation, immuno-affinity capture, tag-based affinityAttorney Docket No.: AUMI-004 / 02WO 348385-2089 capture, protein A / G / L affinity capture, enzyme or receptor-ligand affinity capture, lectin-based affinity capture, carbohydrate and sugar-based affinity capture, capture by carbohydrate-binding modules, synthetic glycopolymer affinity capture, hybridization-based capture, aptamer-based affinity capture, affinity capture based on peptide nucleic acid probes, affinity capture based on protein-nucleic acid interactions, magnetic bead capture, ultrasonic capture, size exclusion chromatography, ion exchange chromatography, hydrophobic interaction chromatography, electrophoresis, dialysis, flow cytometry, field-flow fractionation, AC electrokinetics, and embedding.
33. The method as in any preceding claim, further comprising processing the biological fluid using a microfluidic chip comprising an inlet reservoir, at least one channel, at least one obstruction configured to restrict fluid flow, and an outlet reservoir.
34. The method as in any preceding claim, further comprising processing the separated ECMBs by applying to the separated ECMBs one or more of a protease inhibitor, a histochemical stain, a hematoxylin stain, an eosin stain, a histology stain, an alcian blue stain, a picrosirius red stain, an immunohistochemical (IHC) stain, a multiplex IHC stain, multi-spectral imaging, a protein stain, a nucleic acid stain, and a chemical fixation.
35. The method as in any preceding claim, wherein the ECMB data comprises one or more biomarkers selected from the group consisting of Fibronectin, Tetranectin, Thrombospondin, Galectin-3 Binding Protein (3BP), Talin, Rab27B, Zyxin, DAPI, CD5L, Afamin, Carbonic Anhydrase (CA1), INF2, Clusterin, Victronectin, Gelsolin, S100A9, CD5L, and combinations thereof.
36. The method as in any preceding claim, wherein the cancer parameter comprises one or more of a pre-cancer, a cancer type, and a cancer stage.
37. The method as in any preceding claim, wherein the pre-cancer comprises one or more of pre-cancerous colon or rectal polyps, pre-cancer lesions, and advanced adenoma.Attorney Docket No.: AUMI-004 / 02WO 348385-2089 38. The method as in any preceding claim, wherein the cancer type comprises one or more of non-cancer, colorectal cancer, lung cancer, pancreatic cancer, prostate cancer, breast cancer, esophageal cancer, liver cancer, ovarian cancer, kidney cancer, melanoma, and gastric cancer.
39. The method as in any preceding claim, wherein the ECMB data comprises one more of a flake, a punctum, a decorator, and combinations thereof.
40. The method as in any preceding claim, wherein the flake comprises an eosinophilic structure having an area of between about 5 μm2and about 450 μm2, wherein the punctum comprises a circular structure having an area less than about 15 μm2, and wherein the decorator comprises a circular structure having an area less than about 7 μm2.
41. The method as in any preceding claim, further comprising updating an electronic medical record based on the cancer parameter prediction.
42. The method as in any preceding claim, further comprising generating a report for one or more stakeholders, the report comprising the cancer parameter prediction.
43. The method as in any preceding claim, wherein the subject is a human subject.
44. The method as in any preceding claim, wherein the subject is a non-human subject.