Systems, methods and testing apparatus for assessing platelet samples using machine learning for cancer estimation

A machine learning model with a stacked ensemble architecture and XGBoost feature selection effectively addresses the challenges of identifying biomarkers in TEPs, achieving high accuracy and specificity in early-stage cancer detection by reducing dimensionality and mitigating overfitting, enabling reliable cancer diagnosis and monitoring.

WO2026156436A1PCT designated stage Publication Date: 2026-07-30COPOLY AI INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
COPOLY AI INC
Filing Date
2025-12-30
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Existing cancer diagnostic methods, particularly those using tumor-educated platelets (TEPs), face challenges in identifying robust biomarkers and distinguishing cancer signals from non-cancer signals, leading to inaccurate results, especially in early stages, and are prone to overfitting due to high-dimensional data.

Method used

A machine learning model with a stacked ensemble architecture, utilizing XGBoost for feature selection and a neural network meta-model, is trained to analyze TEPs, reducing dimensionality and mitigating overfitting, while maintaining model diversity and generalizability, and includes a CI/CD pipeline for continuous monitoring and retraining.

Benefits of technology

The system achieves high diagnostic accuracy (>96.47%) and specificity (97.4%) in detecting early-stage lung cancer, with improved computational efficiency and interpretability, and can monitor biomarker changes over time, providing reliable cancer detection and staging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CA2025051768_30072026_PF_FP_ABST
    Figure CA2025051768_30072026_PF_FP_ABST
Patent Text Reader

Abstract

An artificial intelligence-based method for detecting a disease comprises a meta- classification model comprising a plurality of base classification models each generating a respective output and wherein each of such outputs is input to a meta-model, the meta-model comprising an artificial neural network. Variants are described in respect of the meta-model stacked ensemble pipeline, dimensionality / complexity control, and specific architectures designed for tracking disease trajectory through identifying gene pathway crosstalk. The method is particularly suitable for detecting cancer in blood samples comprising tumor educated platelets.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS, METHODS AND TESTING APPARATUS FOR ASSESSING PLATELET SAMPLES USING MACHINE LEARNING FOR CANCER ESTIMATION CROSS REFERENCE TO PRIOR APPLICATIONS

[0001] This application claims the benefit of US Provisional Application No. 63 / 749,134 filed January 24, 2025 and US Provisional Application No. 63 / 863,129 filed August 13, 2025, which are incorporated herein by reference in their entirety.FIELD

[0002] The present description relates to systems, methods and testing apparatus for assessing physical platelet samples extracted from tumor-educated platelets using machine learning. More specifically, the description relates to a computerized testing approach and physical testing apparatuses that are adapted to analyze features of tumor-educated platelets to generate a computer output indicative of whether a proposed improved machine learning model estimates that there is a likelihood of cancer in the features of the tumor-educated platelets greater than a pre-defined threshold. A specific machine learning architecture using a stacked ensemble is described in some embodiments, the architecture addressing technical challenges associated with practical limitations to computing resources.INTRODUCTION

[0003] Cancer remains one of the leading causes of mortality worldwide, with over 19 million new cases and nearly 10 million deaths annually. Early detection of cancer is critical in improving survival rates and reducing the burden of this disease, as treatment outcomes are significantly more favorable when cancer is diagnosed in its early stages. However, traditional diagnostic methods, such as imaging and biopsies, often fail to detect cancer until it has progressed to more advanced stages, thus limiting treatment options and reducing the likelihood of successful intervention. Indeed, diagnosis at stage lll-IV dramatically reduces 5-yr survival rates across many cancers.

[0004] In addition to the traditional methods mentioned, various computer-based diagnostic methods have been proposed that comprise assessment of biomarkers in a patient sample using classification models. Examples of such methods are provided in US 2024 / 0087754, US 2023 / 0263477, US 2020 / 0005901, US 2019 / 360051, US 11,948,684, and WO 2012 / 150275.

[0005] Additionally, one of the issues associated with known diagnostic methods, particularly for cancer diagnostics, relates to the need for obtaining a biological sample for conducting the analysis. For example, with image-based methods and systems, invasive procedures such as tissue biopsies are often required. In contrast, blood-based diagnostic methods have been recognized as a promising option owing to the less invasive blood collection procedure that is required for acquiring the patient sample. In some of these methods, disease detection is based on identifying circulating tumor cells (CTCs) and epigenetic or methylation patterns.

[0006] While these markers are prevalent at late stages, studies have shown that the concentrations of these markers in earlier stages (stages l-ll) are either too low to detect or are indistinguishable from baseline or non-cancer conditions. As an alternative, tumor-educated platelets (TEPs) obtained from blood samples have been proposed as a source of biomarkers for cancer detection. Examples of TEPs in this manner is provided, for example, in IN 2020 / 11042049 and US 2019 / 360051.

[0007] With blood-based methods, even those employing TEPs, a major challenge in developing an accurate computer-based diagnosis platform lies in identifying a robust set of biomarkers or a biomarker panel, and strategies to efficiently elucidate cancer signals from noncancer signals that are inherent to platelet biology, allowing such detection mechanisms to generate accurate results.SUMMARY OF THE DESCRIPTION

[0008] An improved testing approach and physical testing apparatuses are proposed that are adapted to analyze features of tumor-educated platelets (TEPs). As described herein in more detail, a proposed improved machine learning model having specific computing architecture is trained and operated in inference mode to computationally generate estimates whether there is a likelihood of cancer in the features of the tumor-educated platelets greater than a pre-defined threshold.

[0009] The physical testing apparatus can include a trained machine learning model that has been trained using a cohort of samples obtained for a particular demographic. They physical testing apparatus can include, cartridges, reader devices, sequencing kits and instruments, assay reagents, finger prick device, micro-sampling device, and syringes. These physical testing apparatus cooperate, including all combinations and permutations of the listed exemplary apparatuses, to provide a desired sample of the tumor educated platelets, extracted and / or amplified RNA, or RNA sequence read organized data. In a practical implementation, a test method can include using both a proposed wet lab apparatus and proposed dry lab apparatus in accordance with embodiments described herein. Physical samples are obtained or extracted to obtain tumor educated platelets, from which RNA is extracted and sequenced using a physical sequencer instrument or by chemical reagent assays. The sequences and physical sample characteristics are encoded in the form of data objects and data structures. A physical computing apparatus, which may be a special purpose machine, that is configured specifically for maintaining an innovative machine learning model data architecture having a stacked ensemble and trained in accordance with specialized training approaches as described herein, such as training datasets that represent clinical classifications and endpoints. The trained machine learning model is a proposed custom neural network model that includes specific architecture, such as a sequential model and utilizes a stacking ensemble of a multitude of diverse base models that are operated together. In a further variation, the neural network modelis also configured to operate for specific user to monitor biomarker changes over time and / or following medical interventions or medical events. The neural network model outputs a decision support interface that presents a visual graphic representation of a patient’s biomarker panel, including for example a time based dynamic graphic representation of a patient’s biomarker panel with embedded markers indicating medical interventions or medical events. The neural network model outputs a decision support interface that automatically lists available list of recommended treatment regimen candidates. The neural network model outputs structured data or metadata for embedding or appending to patient health records.

[0010] The trained machine learning model is operated as a meta-model that utilizes the predictions of multiple underlying models as input features, and the benefit of this approach includes enhanced technical performance, as well as a reduced propensity for overfitting and a broad mechanism for pattern detection. Given the high dimensionality of the dataset, Applicants are acutely aware of the challenges posed by such a large feature space. High-dimensional data often suffers from the curse of dimensionality, where the volume of the space increases so rapidly that the data becomes sparse, degrading the effectiveness of distance metrics and impairing the model's ability to generalize.

[0011] This sparsity is particularly detrimental to complex models such as neural networks, which require dense, informative representations to learn meaningful patterns. In such settings, models are more prone to overfitting, capturing noise and spurious correlations rather than robust, learnable signals. These are non-trivial technical problems, and the approaches proposed herein are adapted to address these problems.

[0012] To address this issue, Applicants propose an applied feature selection strategy aimed at reducing dimensionality while retaining the most predictive genes, and this strategy is physically implemented through specially configured computer approaches. This step was especially important given the use of a stacking ensemble, which combines several models — including tree-based learners and a neural network — as part of the meta-learning pipeline. While stacking enhances overall predictive power by leveraging the strengths of different algorithms, it also introduces the risk that high-dimensional input data can overwhelm one or more base learners, particularly the neural network. Reducing dimensionality helps maintain model diversity and improves ensemble robustness by ensuring that each model learns from a more informative and compact feature set.

[0013] For the feature selection process, the proposed approach employs the SelectFromModel utility from the sklearn.feature_selection module, using an XGBoost classifier as the base estimator. XGBoost was chosen for its ability to handle high-dimensional, sparse data efficiently and for its robust internal feature importance scoring. The process involved training XGBoost on the full dataset to compute feature importance values, applying a defined importance threshold, and selecting only those genes exceeding the threshold. This approacheffectively ranks features by their contribution to model performance and discards those deemed redundant or uninformative.

[0014] Through this method, Applicants identified a subset of 605 genes that most significantly contributed to the classification of lung cancer versus other cancer types and healthy controls. This reduction not only improved computational efficiency and training time but also enhanced model interpretability, a crucial factor when working in genomics, where the biological significance of selected genes can provide insights into disease mechanisms.Furthermore, by presenting a more focused feature set to the neural network within the stacking ensemble, the system mitigated overfitting risks and promoted better generalization to unseen data.

[0015] Overall, dimensionality reduction via model-based feature selection proved critical in building a performant and interpretable ensemble model. It enabled the proposed approach to balance predictive accuracy with model transparency while ensuring that complex learners, like neural networks, operated within an optimized, information-rich space that fostered learning rather than memorization.

[0016] The trained machine learning model is trained to optimize a loss function represented through its interconnected weights, filters, and parameters, and this loss function can include a binary cross-entropy loss. Different variations of the trained machine learning model can be used for different types of training, such as training using data sets for early-stage cases, as well as for late-stage cases. Once trained, the trained machine learning model can be deployed as a trained model for inference usage. The proposed Oncosage model is designed with a robust infrastructure that allows for continuous monitoring, versioning, and retraining. Applicants have implemented performance tracking mechanisms to detect model drift and data drift over time. This includes monitoring key metrics such as accuracy, precision, recall, and calibration across different subpopulations and time windows, allowing Applicants to identify when the model's predictions may no longer align well with real-world outcomes.

[0017] To support this, the system has a CI / CD (Continuous Integration I Continuous Deployment) pipeline tailored for machine learning models. This pipeline includes automated model versioning, reproducible training runs, and a structured promotion workflow for moving models from development to staging and production environments. The system also use tools such as MLflow and DVC (Data Version Control) to ensure that both the model artifacts and the underlying data and features are version-controlled and auditable.

[0018] Currently, the system is configured to have a retraining pipeline in place that is triggered as the system gathers data from broader demographics. This pipeline incorporates data validation steps, updated preprocessing, and hyperparameter tuning to optimize model performance with each iteration. In a further variation, the system can be configured to refine itself in real time based on incoming measurements, and the retraining loop ensures that as newlabeled data becomes available — especially across cancer subtypes or evolving patient cohorts — the model can be updated and redeployed.

[0019] This infrastructure ensures that Oncosage’s predictions remain accurate, reliable, and aligned with a commitment to clinical utility and real-world applicability.

[0020] The trained machine learning model can be represented as a static set of interconnected weights, filters, and parameters, and a new unknown sample can be provided to the trained machine learning model for processing to generate a predictive output representative of a logit indicating a normalized probability representing the classification output of the trained machine learning model, and this normalized probability can be compared against a threshold value to generate one or more quantized outputs of whether the new sample potentially has no cancer estimated, early stage estimated, or late stage cancer estimated.

[0021] The trained machine learning model used in Oncosage retains the full ensemble structure during inference, including all base models and the meta-model. The system is configured to use a stacking ensemble approach, where multiple base learners make individual predictions, and a meta-model combines these predictions to produce the final output.

[0022] During inference, all base models are executed in parallel to generate their respective outputs. These outputs are then fed into the meta-model, which synthesizes them into a single, robust prediction. This design allows the system to leverage the strengths of different model architectures — some may capture linear patterns well, while others are better at modeling nonlinear relationships or interactions among gene features.

[0023] Unlike some compressed or distilled models, the system deliberately retains the full ensemble at inference time to preserve accuracy and generalizability. While this may introduce a slightly higher computational cost compared to using a single compressed model, it ensures that the predictive performance observed during validation is maintained in production use.

[0024] In summary, all components of the ensemble — base models and the meta-model — actively participate in the inference process, ensuring that Oncosage’s predictions remain both robust and reliable across diverse patient profiles.

[0025] From a practical implementation perspective, the physical testing apparatus can be implemented in the form of an oncology sample testing kit, where a static deployment of the machine learning model is maintained. In some embodiments, the oncology sample testing kit can include a receiver front-end that is configured to receive, a blood sample or isolated platelet samples. Where a blood sample is received, the oncology sample testing kit can include centrifuge mechanisms that are adapted to separate platelet-rich-plasma, and ultimately platelet pellets. The platelets can be provided in the form of pellets that are frozen for RNA-sequencing.

[0026] Lab testing kit: A lab testing kit designed for a clinic or hospital use would contain components that are needed to collect a sample from one patient and send it to a processing lab. Specifically, it will include these components:

[0027] 1) Collection tube: a modified 10 mL BD Vacutainer that is pre-coated with a proprietary mix of RNA-stabilizing additives that are optimally designed for platelets. This would allow same-day or next-day processing without any RNA degradation. As an alternative, one could use an existing 10 mL BD Vacutainer (e.g. Lavender top tubes with precoated K2EDTA which was is used now for blood sample collection). 2) Instruction manual: instructions on how to collect the blood, the volume to collect, how to pack it and send it to a processing lab, instructions on storage conditions, description of the contents of kit, information on online documentation and videos, contacts for customer support. 3) QR code: The tube would be affixed with a QR code that can be scanned into the LIMS of the clinic or hospital, and the OncoSage online portal for sample tracking. 4) Shipping label: the kit would contain a prepaid shipping label and pack for a courier to expedite shipping to the processing lab.

[0028] Once the blood is collected, the customer can arrange for sample pickup with one of the preselected courier partners who will expedite shipping to the closest processing lab. The processing lab will process the blood to extract platelets, extract RNA, and perform sequencing. They can then deposit the sequencing data into the OncoSage platform using the web portal. The data will be used as input in the OncoSage platform, and the output will be the clinical report containing the results of the test (cancer positive, cancer negative, or inconclusive, cancer stage stratification). These results will be sent to the ordering physician and notify the patient. Home testing kit (short term): In an embodiment, the approach is coupled with a physical testing kit designed for home use that contains all of the required materials to collect blood at home. At this stage, the blood collection would still have to be completed by a trained phlebotomist, as such the kit would be designed with that in mind. The contents would include:

[0029] 1) Collection tube: same as above, a modified 10 mL BD Vacutainer that is precoated with a mix of RNA-stabilizing additives, or a BD Vacutainer precoated with K2EDTA. 2) Blood collection service voucher: the kit would include a prepaid voucher that allows the customer to call a telephone number or visit a website to schedule an at-home blood collection appointment with a partner mobile clinic. This service will send a phlebotomist who is trained on the kit to the customer’s home to collect their blood. 3) Instruction manual: instructions for the customer on how to schedule an appointment for blood collection, instructions on how to collect the blood (for the phlebotomist), the volume to collect, how to pack it and send it to a processing lab, instructions on storage conditions, description of the contents of kit, information on online documentation and videos, contacts for customer support. 4) QR code: The tube would be affixed with a QR code that can be scanned into the LIMS of the clinic or hospital, and the OncoSage online portal for sample tracking. 5) Shipping label: the kit would contain a prepaid shipping label and packaging for expedited shipping to the processing lab. 6) Mobile app: to allow the customer to track their sample, monitor progress, view online instruction manuals and videos, and view results. Once the blood is collected, the customer can arrange for samplepickup with a courier who will expedite shipping to the closest processing lab. The processing lab will process the blood to extract platelets, extract RNA, and perform sequencing. They can then deposit the sequencing data into the OncoSage platform using the web portal. The data will be used as input in the OncoSage platform, and the output will be the clinical report containing the results of the test (cancer positive, cancer negative, or inconclusive, staging stratification). These results would be sent to the ordering physician and notify the patient.

[0030] Home testing kit: A testing kit can be included that is designed for home use that contains all of the required materials for the customer to collect their blood themselves. This would eliminate the requirement for a phlebotomist to visit the customer’s home for blood collection. The contents would include: 1) User-friendly blood collection device: this can include either a finger prick device, or a micro-sampling system such as a Mitra or TAP system to facilitate the customer to collect a small volume of blood. 2) Collection cartridge: A pre-loaded cartridge for the collection system that contains RNA stabilizers to preserve platelet RNA during shipping. The RNA stabilizers will be optimized for downstream transcriptomics use. 3) Instruction manual: instructions on how to collect the blood, how to pack it and send it to a processing lab, instructions on storage conditions, description of the contents of kit, information on online documentation and videos, contacts for customer support, etc. 4) QR code: The tube would be affixed with a QR code that can be scanned into the LIMS of the clinic or hospital, and also the OncoSage online portal for sample tracking. 5) Shipping label and box or envelope: the kit would contain a prepaid shipping label and pack for a courier to expedite shipping to the processing lab. 6) Mobile app: to allow the customer to track their sample, monitor progress, view online instruction manuals and videos, and view results. Once the blood is collected, the customer will place the cartridge in the provided shipping pack with a prepaid shipping label and arrange for sample pickup with one of the preselected courier partners who will expedite shipping to the closest processing lab. The processing lab will process the blood to extract platelets, extract RNA, and perform sequencing. They can then deposit the sequencing data into the OncoSage platform using the web portal. The data will be used as input in the OncoSage platform, and the output would be the clinical report containing the results of the test (cancer positive, cancer negative, or inconclusive). These results would be sent to the ordering physician and notify the patient.

[0031] The oncology sample testing kit extracts the RNA-sequencing and converts the RNA sequence into vectorized data inputs that are then provided to the trained machine learning model for operation in inference mode. Operation in inference mode includes a computer processor and coupled memory maintaining the trained machine learning model operating in conjunction with a data receiver that is configured to receive the vectorized data inputs and to generate a prediction of whether the vectorized data inputs are classified into one of the classes based on a corresponding probability representing logit generated by passing the vectorizeddata inputs into the classifier. The operation in inference mode can be conducted over a period of time such that periodic monitoring of progressive samples can be measured, and the corresponding changes in classifier outputs can be tracked in data storage. For a particular user, the corresponding changes in classifier outputs are used to determine a rate of change and increase in classification probability I outputs for each particular classification, and patterns in change can also be used to update an overall classification.

[0032] In a further embodiment, an approach is proposed to track patterns through repeated testing on the system, and the approach includes establishing a data object that represents a digital twin. The platform is configured to establish and provide a snapshot of the gene expression profiles occurring in platelets, and as a proxy, in any tumors that may be present. Platelets naturally replenish in the body approximately every seven days. Therefore, new platelets will always be circulating and provide surveillance of the body, and the approach can include collecting samples at different timepoints in order to monitor changes in a patient for minimum residual disease (MRD) or for longitudinal testing.

[0033] The initial sample, So, taken from a patient, P, will be the baseline timepoint, to. The next subsequent sample taken from P will be S / at timepoint f,. Every subsequent sample will be Si+i at timepoint ti+i. The results from every / 1hsample at every / 1htimepoint for P will be tracked to monitor how cancer signals change overtime. Additionally, complete gene expression profiles for P will also be tracked over the time. OncoSage will create a digital twin of the patient that is essentially their gene expression snapshot at every timepoint, ti. OncoSage’s inference model will then compare which genes and molecular pathways are changing over time and provide this information as a summary in subsequent reports to the ordering physician via the OncoSage portal. OncoSage will also use this information on which genes and pathways are being altered to determine the likelihood of milestone events / processes including but not limited to tumor shrinkage, tumor growth, tumor metastasis, drug resistance, and recurrence. This will be achieved by tracking mutations and changes in gene expression patterns, predicting which pathways are likely to be altered next, and inferring which processes are likely to be affected. This will be accomplished by a separate Al model that can take as input the results at different timepoints and using a foundational large language model trained on nucleotide base letters to predict where changes will occur. This will take into consideration the global network of genes and pathways and the crosstalk between genes and pathways.BRIEF DESCRIPTION OF THE FIGURES

[0034] The features of certain embodiments will become more apparent in the following detailed description in which reference is made to the appended figures wherein:

[0035] FIG. 1 illustrates the steps followed in the training and testing of the method and system according to one embodiment.

[0036] FIG. 2 illustrates an application of the methods described herein for cancer screening.

[0037] FIG. 3 illustrates an application of the methods described herein for screening and / or monitoring at-risk populations for cancer.

[0038] FIG. 4 shows a graph comparing recall performance across individual models and the stacked ensemble model.DETAILED DESCRIPTION

[0039] In one embodiment, there is described herein an artificial intelligence (Al) bloodbased early cancer detection platform designed to identify one or more cancer types with high accuracy, specificity, and sensitivity. As discussed above, conventional diagnostic techniques often rely on invasive procedures or on biomarkers that are only detectable at later cancer stages. The present description instead uses tumor-educated platelets (TEPs) as the biological sample for performing the diagnosis.

[0040] The OncoSage tool can be physically implemented as a compact benchtop diagnostics instrument that is roughly the size of a small desktop printer or an espresso machine. This instrument is built for point-of-care or decentralized lab use (e.g., outpatient clinics, mobile diagnostics, pharmacists, rural clinics, or undeveloped countries). The device could be a closed, cartridge-based system that automates all steps from blood input and platelet isolation to Al-driven cancer detection results. The device would have an integrated workflow that includes four core phases: 1) Blood collection and cartridge loading: a clinician, technician, or patient collects a small volume of blood (2-5 mL) via finger prick or standard venipuncture into a preloaded microfluidic cartridge. The cartridge contains all required buffers, stabilizers, and compartments for processing. The cartridge is inserted into the OncoSage device via a front-facing port. 2) Automated platelet isolation: inside the cartridge, microfluidic separation channels isolate platelets. This is achieved using one or a combination of passive filtration, dielectrophoresis, or mini-centrifugation using a miniature rotor. Platelets are captured in a dedicated chamber within the cartridge without requiring any external equipment or tools. All other components of the blood (e.g., red blood cells, white blood cells) are discarded in dedicated waste chambers. In some embodiments, captured or extracted platelets are further purified, for example, by antibody-based purification, FACS flow cytometry, or using magnetic beads. 3) RNA extraction and amplification: The device will then lyse isolated platelets and automatically extract RNA into a dedicated chamber using integrated magnetic bead or silica column technology. RNA is purified and quality checked in-line using one or a combination of microfluidic-based fluorometry or electrochemical biosensors. The RNA chamber will then perform reverse transcription and pre-amplification using lyophilized reagents stored inside the cartridge. The device will achieve this using an isothermal amplification step, avoiding the need for a full thermocycler and allowing for a smaller footprint. 4) Sequencing and Al analysis: Thedevice includes a low-throughput sequencer. The sequencer may be nanopore-based or semiconductor-based. The sequencer reads transcriptomic data directly from platelet RNA or the cDNA inside the dedicated chamber in the microfluidic cartridge. The OncoSage Al model will be either run directly on the device using onboard semiconductor chips, or the Al model will be accessed through a secure cloud connection to OncoSage servers. The Al model will use the sequencing results as input and detect any cancer signatures. Results will be displayed on a built-in screen and synced securely with the patient’s EHR / EMR platform and the cloud-based OncoSage dashboard.

[0041] This apparatus will utilize disposable cartridges that are preloaded with all the buffers and reagents required for all four phases, including barcodes and sample ID tracking chips. The device will have upgradeable software and AI / ML models that will allow for new cancer types or use cases. The apparatus can identify what type of cartridge is loaded into the device based on two-way communication with the cartridge. TEPs are circulating platelet cells which, directly or indirectly, take up genetic material (DNA or RNA) or tumor cells they encounter in the body. In effect, TEPs essentially build and maintain a catalogue or “memory” of the tumors they encounter in the body. However, while previous studies have postulated using TEPs for cancer diagnosis or detection, and while some studies have demonstrated potential in implementing machine learning or Al methods to identify cancer using TEPs, they used out-of-the-box algorithms and fell short of reliable detection with poor efficacy, sensitivity, or specificity of both late-stage and early-stage. The present inventors have developed a unique non-invasive and cost-effective solution for early cancer screening and diagnosis that is capable of detecting cancer as early as stage I.

[0042] According to one embodiment of the present description, the inventors have surprisingly found that a unique combination of multi-omics analysis and machine learning algorithms performed on blood samples achieves a significant performance improvement over many of the known computer-based cancer diagnosis systems. The inventors have developed and trained the Al model described herein on extensive patient datasets, enabling it to detect subtle gene expression patterns associated with early-stage or late-stage cancer. Preliminary studies by the inventors have demonstrated the platform’s ability to detect lung cancer, with a diagnostic accuracy >96.47%, specificity of 97.4%, and sensitivity of 90.47%.

[0043] The training and inference operations of the model are distinctly separated to support a high-performance, multi-model architecture purpose-built for blood-based cancer diagnostics. During training, the system ingests a high-dimensional dataset derived from RNA expression profiles. Using a model-based feature selection technique — specifically XGBoost within a SelectFromModel framework — the system identifies a reduced subset of 605 genes that exhibit the highest predictive value across early-stage and late-stage cancer cases. This step is critical not only for interpretability and computational efficiency but also to prevent overfitting,particularly given that the final architecture includes a neural network as both a base learner and a meta-learner. By reducing the feature space before modeling, the approach mitigates the effects of the curse of dimensionality, ensuring the neural network learns generalizable representations. Following this preprocessing phase, the core training pipeline starts. The model architecture employs a stacking ensemble strategy, which comprises four diverse base classifiers — XGBoost, LightGBM, Random Forest, and a custom feed-forward neural network — followed by a neural network meta-model trained on their outputs. Each base model is independently trained on the same reduced gene expression dataset. Their output, a probability of lung cancer presence, is then used to train the meta-model, which learns how to optimally combine predictions, correct model-specific biases, and boost final classification performance. Importantly, the base models are not compressed or abstracted away post-training. Their full configuration and learned parameters are preserved, meaning the trained ensemble retains all individual decision pathways. This has a direct effect on the system's ability to achieve high sensitivity and specificity, particularly in early-stage cancer detection, which remains one of the most challenging problems in the field. The meta-model is trained exclusively on out-of-fold predictions from the base learners, ensuring that it never sees data used to train its input predictors. This strict separation is vital to avoid information leakage and ensure that the stacked ensemble generalizes well to unseen samples. The proposed custom neural network meta-model has a minimal architecture — one hidden layer with dropout regularization — to reduce overfitting on top of already-learned predictions. It uses a sigmoid activation function to output a calibrated probability between 0 and 1, representing the likelihood of lung cancer. In inference mode, the architecture remains intact. All trained base models and the meta-model are loaded as part of the ensemble. When a new blood sample is processed and converted into a vector of the 605 gene expression values, it is simultaneously passed through each base model. Their outputs — again, cancer probabilities — are then passed as input to the meta-model, which produces the final diagnostic prediction. No approximation, averaging, or simplification is applied to the ensemble at inference.

[0044] This full-retention strategy is a distinguishing technical feature, as many systems replace the ensemble with a distilled or single-model version for speed or portability. The current invention maintains the integrity of the ensemble because Applicants have observed, empirically, that the diversity of the component models — including tree-based and neural architectures — is critical to achieving accuracy and sensitivity. This separation of concerns — where training involves a coordinated pipeline of feature reduction, independent model optimization, and meta-learning, while inference is a deterministic, multi-model forward pass over the preserved architecture — is central to the reliability of the system. It also enables modular upgrades: as additional data is acquired or the system is expanded to new cancer types, models can be retrained and swapped into the ensemble without disrupting theunderlying structure. This adaptability, together with the use of a stacked ensemble in which both tree-based and neural models are preserved at inference, represents a novel and technically robust implementation that contributes directly to the system’s superior diagnostic performance and distinguishes it from traditional approaches in the field.

[0045] Although the present description is focused on its application for lung cancer detection, i the teaching herein is not limited to lung cancer and can indeed be used for multiple cancer types, including but not limited to breast, bladder, colorectal, esophageal, glioblastoma, head and neck, kidney, liver, ovarian, pancreatic, prostate, and / or lymphomas. Additionally, although in one preferred embodiment, the methods and systems described herein are suited for “early” stage cancer detection, it will be understood that such methods and systems can be used to detect cancer at any stage of disease. Moreover, although the terms “early” and “late” are used, it will be understood that more refined stratification of a disease state is possible.

[0046] Described in more detail below are embodiments that are meant to illustrate the description, in particular, the sample collection and data retrieval / analysis steps involved in obtaining preliminary dataset, the identification of TEP biomarkers, the creation of biomarker panels, the development and training of the various algorithms, and the testing of the system. The description will reference FIG. 1.

[0047] Cancer risk and gene expression can vary significantly by age, sex, ethnicity, smoking history, comorbidities, and family history. This metadata will be integrated with transcriptomics to improve the model accuracy to improve early-stage detection and reduce false positives. This metadata will be used as secondary input into the OncoSage Al model. These features will be passed through embedding layers and fused together with the transcriptomic input data in the model architecture. This will enable risk stratification based on a patient’s demographics (e.g., the same RNA signature may be interpreted differently for a 35-year-old non-smoker vs. a 65-year-old smoker with family history of cancer).

[0048] In this variant approach, this system can flag atypical results (e.g., high-risk RNA signature in a low-risk demographic) for additional follow-up, monitoring, or confirmatory downstream testing. With sufficient real-world data on specific demographics, the proposed system can train subpopulation-specific models to ensure equitable performance across ethnic groups and underserved populations.

[0049] Isolation of tumor-educated platelets (TEPs) and RNA extraction - As shown in FIG.1 at 10, blood was collected from patients using established blood collection procedures. At 12 platelets were isolated from whole blood samples within 48 hours of blood collection, until which whole blood samples were stored at 4 °C. Specifically, 10 mL of blood was collected into EDTA-coated purple-capped BD Vacutainers (BD, #367863). Platelets were isolated from whole blood samples using established two-step centrifugation procedures. First, whole blood samples were centrifuged at 120 g for 20 minutes to separate platelet-rich-plasma (PRP) from other bloodcells. The PRP component was carefully removed (90-95%) to prevent contamination with other blood cells. The PRP was then centrifuged at 360 g for 20 minutes to pellet the platelets. The platelet pellets were then resuspended in cell lysis buffer and frozen at -80 °C. Samples were then thawed in batches on ice and standardized protocols for RNA extraction were followed to extract the RNA for RNA-sequencing (RNA-seq). All centrifugation steps were performed at room temperature.

[0050] RNA preparation and RNA-seq - The RNA sequencing and analysis are indicated at 14 in FIG. 1. RNA samples were quantitated using Qubit™ 2.0 Fluorometer (Thermo Fisher Scientific) and quality control analysis using Agilent 2100 Bioanalyzer™ (Agilent Technologies #G2939BA) and RNA 6000 Nano Kit™ (Agilent Technologies #5067-1511). Ribosomal RNA removal, library preparation, and cDNA from each unique patient was prepared using SMART-seq Total RNA Pico Input kit with UMIs (Takara #634354 and #634752) and labelled with an indexing barcode. The target fragment length for cDNA was >300 bp. High-throughput sequencing was performed on an Illumina NextSeq™ 500 or Hiseq™ 2500 or 4000 platform (single-read, mid-output, 150-cycle kit) and raw FASTQ™ files were generated.

[0051] RNA-seq analyses - A preprocessor was run to prepare the reads from above for further analysis. Specifically, the preprocessor processed all the raw FASTQ™ files to trim adapter sequences and to identify indexing barcodes to de-multiplex individual patient reads. The reads from the FASTQ™ files for each patient were then aligned using the Spliced Transcripts Alignment to a Reference (STAR) software (ver. 2+) to the human genome (Homo sapiens GRCh37 / hg19) to generate read counts. A custom program, DiffEx™, developed for analyzing differential expression, was then run to process the read counts to determine gene expression patterns in each patient sample. DiffEx™ reads in all patient samples and combines the metadata for each patient. Such metadata may include clinical parameters provided by the ordering clinician (such as cohort ID, identification of symptomatic or asymptomatic sample, identification of type of cancer, etc.) and other patient-specific data (such as age, gender, ethnicity, etc.). The present description is not limited to any particular metadata and other relevant data for conducting the analysis will be apparent to persons skilled in the art.

[0052] All reads are incremented by a factor of 1 and the trimmed mean of M-values (TMM) normalization method is performed on each patient sample. Three additional manipulations are then performed: 1) each unknown patient sample is normalized to a cohort of known negative controls (asymptomatic cases); 2) each unknown patient sample is normalized to a group of genes that have been identified to be stable or unchanged between cancer and non-cancer cases; and 3) each unknown patient sample is normalized to the mean read count of each sample. To carry out the process described in (2) above, DiffEx™ utilizes a new custom-developed method called HKG-CLR, which utilizes centered log ratios with housekeeping genes. Briefly, HKG-CLR identifies genes above a minimum expression threshold and below amaximum covariance that are expressed stably across a minimum threshold of patient samples. Once these genes are identified, HKG-CLR then uses those genes as housekeeping genes to normalize each sample without the need for a known control sample. Then, utilizing each of these three “normalized” values for each sample, DiffEx™ computes three distinct Iog2 fold change values for every transcript in each patient sample.

[0053] DiffEx™ is a separate upstream program that was developed to perform upstream preprocessing analysis of the FASTQ files prior to input into the Al component of the OncoSage pipeline. Therefore, DiffEx™ is a part of OncoSage and is used in Step 14 in FIG. 1. The above description captures the major steps of DiffEx™. DiffEx™ is a module Applicants developed for OncoSage to take the read counts generated by the preceding step, by the preprocessor module, and calculate differential gene expression, hence the name DiffEx™. A challenge and limitation of blood testing for cancer via genomic or transcriptomic approaches is that noncancer control samples (asymptomatic cases) may not be readily available. This is normally essential in preclinical and clinical experimental design to interrogate gene expression profiles in tumor samples and allow an algorithm to discern gene signatures that are abnormal.

[0054] To overcome this limitation, Applicants devised and implemented a three-part strategy. First, when asymptomatic cases are available, DiffEx™ will normalize the expression of test samples to those asymptomatic cases. In the event where asymptomatic cases are not available, DiffEx™ will utilize the HKG-CLR method described above to normalize test samples to a collection of stable genes which are not known to significantly change or alter expression between cancer and non-cancer cases. Additionally, in the absence of asymptomatic cases, DiffEx™ will also normalize asymptomatic cases to the mean read count of each sample. This method normalizes out any large variations in read counts in biological samples and technical samples. DiffEx™ can be configured to run any one of these methods or a combination of multiple methods to produce one or more Iog2 fold change values for every transcript in a sample.

[0055] Model development and validation cohort - To develop the Al models and validate the performance of the presently described platform, the system was utilized in relation to a cohort of 7,361 human blood samples containing 5,380 healthy or asymptomatic or non-cancer patients, 990 samples containing patients with other types of cancers, including glioma, pancreatic, breast, head and neck, esophageal and prostate cancers, and 982 lung cancer patients. Following the above methods with this sample set produced a dataset containing 57,823 transcripts detected across the 7,361 patients.

[0056] Al component - After the physical sample collection steps and data preprocessing steps described above, the Applicants obtained an initial dataset, as shown at 16 in FIG. 1, comprising gene expression levels for 57,823 transcripts across 7,361 patient samples. As noted above, this included both lung cancer patients, patients with other types of cancers, andhealthy controls. However, a substantial number of genes exhibited missing values (NaNs, or “not a number” values) in their expression data. To address this issue, the system was configured to perform data cleaning, as shown at 18, by removing genes with incomplete data, thereby ensuring a dataset without missing values. Following the cleaning process, the refined dataset consisted of 7,361 patient samples and 4,443 genes with complete expression profiles.

[0057] To establish a baseline for model performance, Applicants began by utilizing the entire dataset to assess how well different classifiers could predict lung cancer presence. This initial step involved splitting the dataset into two subsets, wherein 80% was used for training and 20% was used for testing. The training set was used to train the models, allowing them to learn patterns and relationships within the data. The testing set, comprising unseen data, was used to evaluate the models' performance and generalization capabilities - how well they can make accurate predictions on new, unseen data.

[0058] Preprocessing of training dataset - applying classifier - Given the high dimensionality of the dataset - which consisted of 4,443 genes as features - the Applicants faced challenges associated with large feature spaces. High-dimensional data can lead to the “curse of dimensionality”, where the feature space becomes so vast that the available data becomes sparse. This sparsity can make it difficult for models to learn meaningful patterns, potentially reducing their predictive performance and increasing the risk of overfitting. Overfitting occurs when a model learns the noise in the training data to the extent that it negatively impacts its performance on new data. To explore different modeling approaches and mitigate these challenges, the Applicants selected four classifiers known for their ability to handle complex datasets: Decision Tree, Random Forest, XGBoost, and LightGBM.

[0059] 1) Decision Tree: This is a non-parametric supervised learning algorithm used for classification and regression tasks. It works by partitioning the data into subsets based on the value of input features, forming a tree-like structure of decisions. The benefits of decision trees include their simplicity and interpretability. They can handle both numerical and categorical data and require little data preprocessing. However, they are prone to overfitting, especially in highdimensional spaces, because they can create overly complex trees that capture noise in the data. 2) Random Forest: An ensemble learning method that builds multiple decision trees and merges their results to improve predictive accuracy and control overfitting. Each tree in a random forest is built from a random subset of features and data samples, which introduces diversity among the trees. The aggregation of multiple trees helps to reduce variance and improve generalization to unseen data. Random forests are robust to overfitting and can handle high-dimensional data better than individual decision trees. 3) XGBoost (Extreme Gradient Boosting): A scalable and efficient implementation of gradient boosting algorithms. XGBoost builds additive models in a forward stage-wise fashion; it allows for the optimization of arbitrary differentiable loss functions. The algorithm focuses on speed and performance, handlingmissing values internally and providing regularization parameters to prevent overfitting.XG Boost is known for its high predictive accuracy and has been a popular choice in machine learning competitions. 4) LightGBM (Light Gradient Boosting Machine): Another gradient boosting framework that uses tree-based learning algorithms, designed to be highly efficient and support parallel and GPU learning. LightGBM introduces novel techniques like Gradient-based One-Side Sampling (GOSS) and Exclusive Feature Bundling (EFB) to reduce the number of data instances and features, respectively. This makes it faster and more memory-efficient, particularly beneficial when dealing with high-dimensional data. LightGBM tends to perform well with large datasets and offers faster training speeds compared to other boosting algorithms.

[0060] By applying these classifiers, the Applicants aimed to understand how different algorithms perform in the context of the high-dimensional genomic dataset. Each model has unique advantages that could help in handling the complexities of the data, as outlined below.

[0061] 1) Decision Trees provide a straightforward interpretation but may struggle with overfitting. 2) Random Forests offer robustness against overfitting and can manage highdimensional data more effectively. 3) XGBoost delivers high accuracy and includes mechanisms to handle missing data and prevent overfitting. 4) LightGBM excels in speed and efficiency, making it suitable for large-scale and high-dimensional datasets.

[0062] The results obtained from executing the above-mentioned classifiers are summarized in Table 1. These results provided insights into the predictive capabilities of each model and guided the further optimization steps.

[0063] Table 1: Results from Classifiers

[0064] Accuracy: measures the proportion of correct predictions made by the model out of all predictions. Precision: is the ratio of true positive predictions to the total number of positive predictions made by the model. Recall: measures the ratio of true positive predictions to the actual number of positive cases. Although the initial results were promising, the utilization of over 4,000 variables (genes) in the predictive models presented challenges inherent to highdimensional datasets. High dimensionality can lead to several issues, such as increased computational complexity, difficulty in model interpretation, and a higher risk of overfitting, where the model performs well on training data but poorly on unseen data due to capturing noise instead of underlying patterns.

[0065] To address these concerns and enhance the model's interpretability, the Applicants aimed to identify the most influential genes contributing to the prediction of lung cancer. After comparing the performance of the four classifiers noted above, Applicants selected XGBoost as the primary model due to its superior performance metrics. XGBoost is a scalable and efficient implementation of gradient boosting algorithms, known for its speed and performance, especially with large-scale and high-dimensional data. It offers several advantages.Regularization: Incorporates L1 (“Lasso” Regression) and L2 (Ridge Regression) regularization to prevent overfitting. Handling Missing Values: Automatically learns the best direction to handle missing values. Parallel Processing: Supports parallel tree construction, speeding up computation. Flexibility: Allows optimization of arbitrary differentiable loss functions.

[0066] Preprocessing of training dataset - applying feature selection - With XGBoost as the chosen model, the next step was to perform feature selection, as shown at 20 in Fig. 1, to reduce the number of genes to a subset that significantly contributes to the model's predictions. Feature selection is crucial in high-dimensional data for several reasons: Improved Model Performance: Reduces overfitting by eliminating irrelevant or redundant features. Reduced Computational Cost: Decreases training time and resource consumption. Enhanced Interpretability: Simplifies the model, making it easier to understand and interpret the results. Focus on Significant Features: Helps in identifying genes that may have biological significance in lung cancer.

[0067] The system was configured to employ the SelectFromModel method from the scikit-learn™ library's sklearn.feature_selection module. The SelectFromModel is a meta-transformer that uses a base estimator (in this case, the XGBoost classifier) to determine the importance of each feature. It selects features based on important weights provided by the model. These were obtained using the following steps: 1) Training the Base Estimator: The system trained the XGBoost model on the dataset to obtain feature importances. 2) Thresholding: A threshold value was implemented so as to select only those features with important weights above the threshold. 3) Feature Selection: The method transforms the dataset by selecting only the features deemed important.

[0068] By applying SelectFromModel with XGBoost, the system effectively ranked the genes in the dataset according to their contribution to the model's predictive capability. This process identified 605 genes as the most significant features for distinguishing between lung cancer patients and healthy controls. The reduction from over 4,000 genes to 605 not only streamlined the model but also enhanced its performance by eliminating noise from less informative genes. This refined model is therefore more computationally efficient and has improved generalization capabilities on unseen data. Furthermore, the identification of these 605 genes provides valuable insights into the genetic factors associated with lung cancer. These genes may warrant further biological investigation to understand their roles in cancerdevelopment and progression, potentially contributing to advancements in diagnostics and therapeutics. In summary, by selecting XGBoost for its superior handling of high-dimensional data and implementing feature selection through SelectFromModel, Applicants improved the model's performance and interpretability. This approach allowed the system to be tuned to focus on the most impactful genes, ultimately enhancing the system’s ability to predict lung cancer accurately.

[0069] Final Model - After identifying the most significant 605 genes through feature selection process, to the next step was developing an optimal machine learning model to predict the presence of lung cancer. To enhance predictive performance and leverage the strengths of different algorithms, Applicants configured the system to implement a “stacking ensemble”, as shown at 22 in FIG. 1. This method combines the predictions of multiple machine learning models and uses a final estimator, known as a meta-model, to make the ultimate prediction.

[0070] Stacking, or stacked generalization, is an ensemble learning technique that blends the predictions of several base models to produce a new set of predictions. These predictions are then used as inputs for a higher-level model, namely, the meta-model. The purpose of this methodology is to harness the diverse strengths of various algorithms, capturing a wider array of patterns in the data than any single model could achieve alone. Put another way, base models (or, Level-0 Models) are the initial machine learning models that make predictions based on the training data, whereas a meta-model (or Level-1 Model) takes the outputs of the base models as inputs and learns to make the final prediction. The meta-model thus corrects the biases or errors of the base models.

[0071] The stacking approach offered a few technical advantages, some of which are as follows. 1) Enhanced Performance: By combining different models, stacking can improve predictive accuracy beyond what individual models can achieve. 2) Reduced Overfitting: The meta-model can learn to correct the biases of the base models, leading to better generalization on unseen data. 3) Diversity of Models: Utilizing models with different underlying algorithms captures a broader spectrum of data patterns. Each model may pick up on different aspects of the data, and the meta-model can integrate these insights.

[0072] In a proposed stacking ensemble of some embodiments, the approach included four diverse base models (as shown at 24 in FIG. 1): The stacking ensemble architecture developed represents a highly specialized and original machine learning configuration, specifically engineered for blood-based lung cancer detection using gene expression data. Beyond the use of multiple models in ensemble learning, the proposed approach introduces a series of structural, functional, and interoperability innovations that collectively distinguish it from alternate implementations. Importantly, although the approach employs four base models, one is a neural network — a technical choice that enhances both the diversity and complementarity of learning strategies within the ensemble, and forms the foundation for the proposed novel meta-learning configuration. At the base level, a diverse set of classifiers are used: XGBoost, LightGBM, Random Forest, and a custom neural network. These models are selected not just for their individual performance but for their heterogeneous learning mechanisms. XGBoost and LightGBM are gradient boosting algorithms that excel in handling structured, high-dimensional data with missing values and non-linear boundaries. Random Forest introduces variance reduction through bagging, while also being robust to outliers and noise. The fourth model, a neural network, is explicitly designed to capture nonlinear gene-gene interactions that treebased models are typically unable to learn. This heterogeneity is critical: each model builds a fundamentally different representation of the same 605-gene feature space, and this diversity is essential to the function of the meta-model. What sets the proposed architecture apart is the design and behavior of the meta-model. Rather than employing a logistic regression or linear combiner, a custom neural network is proposed that includes a meta-model trained on the soft probabilistic outputs (rather than hard class labels) of the four base models. This neural meta-model learns not to average, but to contextually weigh and recalibrate base model predictions using nonlinear transformations. The use of dropout layers within the meta-model architecture provides not only regularization but also an internal simulation of missing or unreliable base model outputs. This enables the system to remain robust even if one or more base models deliver weak or noisy predictions in real-world clinical data, which is a technical problem that occurs in genomics. Moreover, the meta-model is trained exclusively on out-of-fold predictions from the base models, ensuring zero data leakage and preserving strict boundaries between training and evaluation data. This architecture encourages the meta-model to focus on resolving disagreement between base learners — a capability that is particularly useful when operating in ambiguous regions of the feature space, such as samples with borderline gene expression profiles or conflicting subtype markers.

[0073] A key point of novelty is that only one of the five models in the full stack is a neural network — a conscious departure from trends in deep ensemble learning, which can rely entirely on neural models. By combining ensemble methods (boosting and bagging trees) with a neural meta-learner, Applicants demonstrate that high performance does not depend on brute-force deep learning alone, but on a carefully engineered interplay of model types, each solving a different part of the pattern recognition problem. This strategic inclusion of just a single neural network among the base models, and its role in handling particular interaction effects, helps reduce overfitting, lower computational burden, and maintain interpretability.

[0074] Another important aspect is how the approach manages and aligns the prediction space. All base models are aligned to output calibrated probabilities rather than class labels. These probabilities form a low-dimensional (4D) representation space that is uniquely structured: each dimension is semantically meaningful (i.e. , one model’s probability for cancer), and the meta-model is trained to learn interactions among these outputs rather than directlyfrom the genomic data. This separation of feature space and decision space enables the metamodel to capture higher-order patterns of agreement, divergence, or conditional reliability among the base models — essentially learning when to trust which model, and how much.

[0075] Finally, the ensemble is deployed in a fully preserved form during inference. There is no compression, distillation, or simplification of the architecture. Each base model is executed in parallel at inference time, and their outputs are passed through the same trained meta-model as in training. This guarantees that the behavior observed during validation is retained in production, a critical consideration for diagnostic reliability.

[0076] In sum, the novelty of the approach lies is beyond the use of four models or in combining them, but in the specific configuration: the purposeful inclusion of a single neural base model within a mostly tree-based stack, the use of a custom neural meta-model trained on soft prediction space, and the robustness engineered into the ensemble via dropout, strict out-of-fold separation, and non-linear weighting.

[0077] 1) Random Forest: An ensemble of decision trees that reduces overfitting by averaging multiple decision trees, thereby improving predictive accuracy. 2) XGBoost (Extreme Gradient Boosting, or XGB): An efficient and scalable implementation of gradient boosting algorithms that excels in handling structured data and capturing complex patterns. 3) LightGBM (Light Gradient Boosting Machine, or LGBM): A gradient boosting framework that uses treebased learning algorithms, optimized for speed and high performance, especially with large datasets. 4) Custom Neural Network (CNN): A deep learning model that was developed for modeling intricate non-linear relationships within the data. Further details of the CNN (i.e. , an artificial neural network) are provided in following sections of the present description.

[0078] Custom Neural Network Details - Applicants designed the custom neural network (CNN) as a sequential model built using the Keras™ API in TensorFlow™ and was designed to effectively process the high-dimensional gene expression data. The architecture of the CNN is as follows: 1) Input Layer: Accepts the input features, namely, the expression levels of the 605 selected genes. 2) Hidden Layers: a) First Dense Layer: Contains 64 neurons with the ReLU (Rectified Linear Unit) activation function, as described above, b) Second Dense Layer:Contains 32 neurons, also using the ReLU activation function. 3) Output Layer: A single neuron with a sigmoid activation function, producing an output between 0 and 1.

[0079] Compilation Details of CNN - Loss Function: the system was configured to use Binary Cross-Entropy as a loss function, as described above. Optimizer: for the optimizer, the system was configured to use Adam (Adaptive Moment Estimation), as described above.Evaluation Metric: Applicants used Accuracy, which assesses the proportion of correctpredictions made by the model. It is calculated as the number of correct predictions divided by the total number of predictions.

[0080] Training Parameters of CNN - Epochs: 100. An epoch is one complete pass through the entire training dataset. Training for multiple epochs allows the model to learn and refine its weights over several iterations. Batch Size: 32. This refers to the number of samples processed before the model's internal parameters are updated. Smaller batch sizes can lead to more stable and thorough learning but may increase training time.

[0081] Base Models Configuration - The base models in the stacking ensemble, see 24 in FIG. 1, were configured as follows: XGBoost Classifier: Configured for binary logistic regression with log loss as the evaluation metric. This means it is set up to predict probabilities for binary outcomes and uses logarithmic loss to measure prediction errors. LightGBM Classifier: Utilizes parameters optimized for performance, making it efficient for large datasets. Random Forest Classifier: Consists of 100 decision trees (n_estimators=100) with a fixed random state (random_state=42) for reproducibility. Custom Neural Network: As described above.

[0082] Data Splitting and Training - Early-Stage Lung Cancer Data - The dataset included cancer stage information for a subset of 521 patients, which were categorized as follows: 33 patients with early-stage lung cancer and 488 patients with late-stage lung cancer. Almost half of patients had unknown cancer stages. Of patients with a known stage, approximately 6% were early-stage and the remainder were late-stage. This significant imbalance, with far fewer early-stage, presented a challenge for training a model capable of detecting lung cancer in its initial stages. To address this imbalance and ensure the model could learn patterns associated with early-stage lung cancer, Applicants modified the data splitting strategy as follows:

[0083] a) Stratified Sampling Based on Cancer Stage: Applicants performed an 80 / 20 split of the dataset consisting of 7361 patients for training and validation, respectively, but with stratification on cancer stage for the subset of patients where this information was available. In particular, Applicants used a Training Set, as shown at 28 in FIG. 1, comprising approximately 80% of the early-stage and late-stage cases and a Validation, or Testing Set, as shown at 30, comprising the remaining approximately 20% of early-stage and late-stage cases, b) Inclusion of Patients with Unknown Stages: Patients in the dataset without stage information were randomly assigned to the training or validation sets, maintaining the overall 80 / 20 split. Since cancer stage was not used as a feature during model training, including these patients did not introduce bias based on or due to stage.

[0084] Purpose and Benefits of Stratification: Stratified sampling techniques are known in the art and the benefits of same are also well known. Some benefits pertinent to the present description are the following. 1) Ensure Representation: Stratifying the data based on cancer stage allowed the proposed approach to have a proportional representation of early-stage cases in both the training and validation sets, despite their limited number. 2) Model Exposure: Thisensured that the model was exposed to early-stage lung cancer patterns during training, which is important for developing a model capable of early detection. 3) Performance Evaluation: Having early-stage cases in the validation set enabled the proposed approach to assess the model's effectiveness in detecting lung cancer at its initial stages, even though stage information was not part of the training features.

[0085] By carefully stratifying the data, the approach enhanced the model's ability to detect lung cancer in early stages, addressing the imbalance in the dataset and ensuring a fair evaluation of the model's performance across different cancer stages.

[0086] Meta-Model Design - For the meta-model, shown at 26 in FIG. 1, the approach proposes a custom neural network (CNN) to take as input the predictions of the base models mentioned above and provide an output. The architecture of the meta-model was designed as follows: 1) Input Layer: Matches the number of base models. This layer accepts the predictions from each of the base models as inputs. 2) Hidden Layers: Comprises a configurable number of dense layers (The system in this example used one hidden layer) with a specified number of neurons (default is 32) and the ReLU activation function. 3) Dropout Layers: These were implemented after each hidden layer to prevent overfitting. The system in this example used a dropout rate of 0.2, meaning 20% of the neurons are randomly dropped during training. 4) Output Layer: This was implemented as a single neuron with a sigmoid activation function for binary classification output.

[0087] Some of the parameters used in developing the meta-model are discussed below. As discussed above, the Adam optimizer was implemented in the meta-model to adjust the learning rate during training for faster convergence. By adapting the learning rate for each parameter, Adam helps the model find the optimal weights more efficiently. The default number of neurons in each hidden layer was 32. As known in the art, neurons are the basic computational units in a neural network. More neurons can allow the model to capture more complex patterns but may increase the risk of overfitting. The default number of hidden layers in the meta-model was 1. Adding layers allows the model to learn more abstract representations of the input data.However, too many layers can make the model overly complex. The dropout rate refers to the fraction of the input units to drop during training (default was 0.2). Dropout helps prevent overfitting by ensuring that the model does not become too reliant on any individual neuron.

[0088] Training the meta-model - The process, shown at 34 in FIG. 1, involved the following steps: 1) Base Model Predictions. Each of the base models was trained on the training set, as shown at 28 in FIG. 1. Thereafter, predictions are made on the validation set, resulting in a new dataset of predictions. 2) Meta-Model Training. The meta-model was trained using the predictions output from the base models as input features. The meta-model thus learns to weigh and combine these predictions to make a final decision. For example, if one base model tendsto perform better on certain types of data, the meta-model can learn to give more weight to its predictions in those cases.

[0089] Testing the meta-model - After training the stacking ensemble model with the selected 605 genes mentioned above, the performance of the meta-model was evaluated (36 in FIG. 1) on the test dataset, 30, described above. The results are provided in Table 2.

[0090] Table 2: Results from Meta-Model Testing

[0091] The detailed evaluation metrics are presented in the confusion matrix and the classification report shown in Tables 3 and 4 below.

[0092] Table 3: Confusion Matrix

[0093] Table 4: Classification Report

[0094] Interpretation of test results - 1. Accuracy - Definition: Accuracy measures the proportion of correct predictions made by the model out of all predictions. Value:0.9647058823529412 (or 96.47%) Implication: The model correctly predicted the lung cancer status for approximately 96 out of every 100 patients in the test dataset. 2. Precision - Definition: Precision is the ratio of true positive predictions to the total number of positive predictions made by the model. It answers the question: "Of all patients predicted to have lung cancer, how many actually have it?" Value: 0.8417721518987342 (or 84%). Implication: When the model predicts a patient has lung cancer, it is correct about 84% of the time. This is important to minimize false positives, reducing unnecessary stress and further testing for patients. 3. Recall (Sensitivity) -Definition: Recall measures the ratio of true positive predictions to the actual number of positive cases. It answers the question: "Of all patients who have lung cancer, how many did the model correctly identify?" Value: 0.9047619047619048 (or 90.47%). Implication: The modelsuccessfully identified approximately 90% of patients who actually have lung cancer. High recall is crucial in medical diagnostics to ensure that few cases are missed.

[0095] Detailed Analysis - Interpretation of confusion matrix - True Positives (TP) = 133: This represents patients with lung cancer correctly identified by the model. True Negatives (TN) = 933: This represents healthy patients correctly identified as not having lung cancer. False Positives (FP) = 25: This represents healthy patients incorrectly predicted to have lung cancer. False Negatives (FN) = 14: This represents patients with lung cancer incorrectly predicted as healthy. As noted above, the metal-model resulted in a low false negative (FN) rate as evidenced by only 14 cases being missed, which is a very important feature in medical settings where missing a diagnosis can have severe consequences. Similarly, the model resulted in a low false positive (FP) rate, where only 25 healthy patients were incorrectly identified as having lung cancer. This unique feature of the meta-model therefore reduces unnecessary follow-up procedures.

[0096] Results for Early and Late-Stage Lung Cancer Detection - To assess the model's capability in predicting lung cancer at different stages, Applicants analyzed its performance separately on early-stage, late-stage, and unknown-stage cases. As discussed above, the dataset included cancer stage information for a subset of patients. The number of cases were as follows: early-stage cases: 33 patients; late-stage cases: 488 patients; and unknown stage cases: 461, representing the remaining patients whose stage information was not available.

[0097] The confusion matrices for the testing on late-stage, early-stage, and unknown stage lung cancer are provided below in Tables 5, 6, and 7, respectively.

[0098] Table 5: Confusion Matrix for Late-Stage Lung Cancer

[0099] Table 6: Confusion Matrix for Early-Stage Lung Cancer

[0100] Table 7: Confusion Matrix for Unknown Stage Lung Cancer

[0101] Discussion of test results for early- and late-stage data testing

[0102] As shown above, by using the meta-model, the number of true positives (TP) for latestage cancer detection was 482 out of 488 (i.e. , patients correctly identified as having late-stage lung cancer) and that for early-stage cancer detection was 29 out of 33 (i.e., patients correctly identified as having early-stage lung cancer). At the same time, the false negatives (FN) were 6 and 4, respectively. In both tests, the true negatives were 0 as no healthy patients were included in the test data. The meta-model achieved 88% accuracy (29 out of 33) in predicting early-stage lung cancer cases in the test set, with only 4 false negatives. Additionally, the model achieved 99% accuracy (482 out of 488) in predicting late-stage cancer cases, with only 6 false negatives. These results demonstrate the unique and unexpected accuracy and reliability realized by the meta-model in not only accurately predicting the presence of lung cancer but also in differentiating between early- and late-stage lung cancer. As discussed earlier in this description, early-stage cancer detection in particular is extremely important in ensuring that adequate medical intervention can be initiated in a timely manner. As a result of the high accuracy achieved and in view of the testing requiring only a blood sample, the model can be integrated into traditional cancer screening programs to easily identify at-risk individuals, especially in populations where early detection rates are currently low.

[0103] Summary of unique features - As will be understood from the embodiments described herein, the methodology developed by the present inventors has several unique features that address complex challenges in cancer detection. Some of these features are summarized: 1) Innovative Use of High-Dimensional Genomic Data - Selective Gene Panel Identification. Starting with over 57,000 genes, the present methodology employed advanced feature selection techniques to identify a critical panel of 605 genes most significant for lung cancer prediction. This precise selection enhances model efficiency and focuses on biologically relevant markers. Effective Handling of Sparse Data. The approach described herein addresses the challenges of missing data (NaNs) and high dimensionality by cleaning the dataset and reducing features without losing critical information, which is essential for processing genomic data effectively. 2) Advanced Stacking Ensemble with Custom Neural Networks - Unique Stacking Ensemble Configuration. The model integrates five diverse base classifiers — XG Boost, SVM, Random Forest, LightGBM, and a custom neural network — into a stacking ensemble. This specific combination is tailored to maximize predictive performance that is particularly suited for cancer detection, such as for lung cancer detection. Distinct Custom Neural Networks used in Base and Meta-Model. The present description incorporates two distinct custom neural networks (CNN) as a component in the base models as well as the meta-model. The CNN in the base model is designed to handle the high-dimensional input of 605 genes by using an optimized architecture with ReLU activation functions and specific layer configurations to capture complex patterns in gene expression data. The CNN in the meta-model uniquely takes the predictions of the base models as inputs and its architecture is specifically crafted tocombine these predictions effectively, using techniques like dropout layers to prevent overfitting. The selection of activation functions (ReLLI and sigmoid), loss functions (binary cross-entropy), optimizers (Adam), and hyperparameters (epochs, batch size, dropout rate) is meticulously tuned to achieve optimal performance, demonstrating a high level of technical refinement. 3) Exceptional Performance in Early-Stage Cancer Detection - As noted above, the model achieves uniquely high accuracy in predicting lung cancers, including both late-stage, and more significantly, early-stage lung cancer cases. This level of performance is rare and addresses a critical need in oncology for early detection methods. By stratifying the dataset to ensure adequate representation of early-stage cases in both training and validation sets, the methodology overcomes the common challenge of imbalanced clinical data. The approach effectively utilizes data from patients with unknown cancer stages to enhance the training process without introducing bias, which is an innovative strategy in medical data analysis. 4) Integration into Clinical Diagnostics - The high accuracy and early detection capabilities of the model make it a valuable tool for integration into clinical workflows, potentially assisting healthcare professionals in diagnosis and treatment planning. In one aspect, once the method of present description identifies the presence of cancer in a subject, a suitable treatment may be administered to the subject. Alternatively, further diagnostic tests may be deemed warranted, such as obtaining and assessing a biopsy to confirm the detection. The methodology also provides a framework that can be adapted to other diseases involving high-dimensional data, highlighting its novelty and utility beyond lung cancer. 5) Addressing Challenges with Unknown Stage Data - The model's ability to maintain high performance despite the presence of patients with unknown cancer stages demonstrates its robustness and adaptability, addressing a common issue in real-world clinical datasets. 6) Consensus-driven triage for complex or ambiguous predictions - There are over 200 different types of cancer which makes training a system that can identify all cancer types challenging due to sample acquisition, cost, and compute resources. This poses a problem for systems that aim to detect cancer from patient biopsies, including blood-based biopsies. For example, the detection system must not confound non-target cancer types or unknown cancer types if a patient presents with such cancers because doing so would lead to false positives and / or misdiagnoses.

[0104] To overcome this challenge, the Applicants have devised an architecture whereby multiple meta-models (as described above) are integrated to achieve a consensus-driven prediction. First, for each target cancer type (e.g. lung cancer), there exists a model that will calculate a prediction for a patient sample. There exists a series of m models for all n target cancer types which are trained on each n target cancer type. Therefore, each sample will be run through all mnmodels to determine outputs, Yn, from each model. These models will determine how closely a patient sample matches each n target cancer type. Secondly, there exists a model, h, that is trained on 1) non-cancer controls composed of asymptomatic (having nounderlying disease or conditions) and symptomatic (having underlying disease or conditions that are not cancer) control patients, 2) all n target cancer types, and 3) additional non-target cancer types. The h model is a large model trained on harmonized, multi-cohort datasets containing both high sample sizes (>200) for controls and n target cancer types and low sample sizes (<200) for non-target cancer types. This enables h to recognize patterns of multiple types of cancer, allowing it to recognize non-target cancer types or unknown cancer types. The model h will process all samples and provide an output X for each sample which states whether the sample is cancer or non-cancer. Finally, a consensus algorithm exists that will observe all outputs, Ynand X for each sample, and use consensus to determine the final reported result for each sample. For the system to conclusively report a particular n target cancer was detected, one and only one of the mnmodels must report a positive output and all other mnmodels must report a negative output (positive = cancer type detected; negative = cancer type not detected) and the model h must report a cancer output. For the system to conclusively report no particular n target cancers were detected, all mnmodels must report a negative output and the model h must report a non-cancer output. If there are any disagreements between mnor h models results, the consensus algorithm will report the result is inconclusive. This architecture reduces confounding bias and reduces false positives and false negatives and also misdiagnoses even for cancer types which the system has not been trained or for cancer types with low sample sizes.

[0105] Discussion - By focusing on the most informative 605 genes and employing a sophisticated stacking ensemble with a custom neural network meta-model, the system is a lung cancer detection platform that provides enhanced predictive capability. The embodiments described herein capitalize on the strengths of various machine learning algorithms and deep learning techniques, resulting in a robust model with improved accuracy and generalization performance. In particular, by effectively combining the strengths of multiple machine learning algorithms through stacking, the model not only attains high overall accuracy but also excels in the critical task of early-stage cancer detection. In other words, although stacking ensembles and neural networks are established methods, the specific combination and configuration used here, in particular the use of custom neural networks in both the base model and the meta-model, is one of the unique features of the approach. This approach not only advances predictive modeling in the context of lung cancer but also provides a framework that can be applied to other high-dimensional biomedical datasets. Furthermore, by identifying key genetic markers through the feature selection protocol mentioned above and the subsequent modeling efforts there is proposed a platform for identifying potential therapeutic targets for treatment of disease, particularly cancer, such as lung cancer.

[0106] The above discussion has focused primarily on the detection of lung cancer. The present description can also be adapted to the detection of other cancers. For example, asnoted above, the dataset used in the development of the model consisted of various cancer types, including lung, breast, esophageal, glioma, head and neck, pancreatic, and prostate. The datasets and models were specifically directed to lung cancer for the purposes of illustrating the efficacy of the proposed approach, particularly for differentiating lung cancer from other cancers. It will be understood that the same approach can be taken for detecting other cancers apart from lung cancer. A lung cancer screening example is illustrated in FIG. 2. As shown, a patient undergoes a routine blood collection, 38, to collect a blood sample. The blood sample is then processed, such as at a processing lab, to isolate platelets from the sample, 40. Various processes and methods may be used for the platelet extraction step and these would be known to persons skilled in the art. Thereafter, RNA is extracted from the platelets and a library is prepared for RNA-sequencing, 42. The step of RNA-sequencing is performed to obtain FASTQ files for all samples. As shown at 42, a preprocessor module is then run to process the FASTQ files into read counts and the DiffEx™ module is run to manipulate the read counts and calculate Iog2 fold change values to determine differential gene expression. The Iog2 fold change values are used as input in the stacking ensemble Al model as discussed above. The stacking ensemble Al model determines if the sample exhibits a gene expression pattern that is indicative of lung cancer. Finally, a report is generated, 46, containing the results for the sample. The results are then provided to the clinician and / or patient.

[0107] The procedure discussed above and illustrated in FIG. 2 can be used for screening cancer in a population of patients. In this example, multiple patients would undergo the same blood collection process as above from which platelets are isolated. RNA from the platelets is then extracted and a library is prepared for RNA-sequencing. The library for each sample is identified, such as with a barcode, with a unique sequence for each patient sample. In this case, multiplexed RNA-sequencing is performed to obtain FASTQ files for all samples. The preprocessor module is run to demultiplex and process FASTQ files into read counts following which, the DiffEx™ module is run to manipulate the read counts and calculate Iog2 fold change values to determine gene expression for all samples. The Iog2 fold change values are used as input in the stacking ensemble Al model and the stacking ensemble Al model determines if the samples exhibit a gene expression pattern that is indicative of cancer. Reports are then generated for each patient sample. An example use case for screening at-risk populations is illustrated in FIG. 3. As shown, patients 48 who meet certain criteria (e.g. at-risk individuals) for cancer testing undergo routine blood collection 50 to collect blood samples. The blood samples are processed and evaluated on the invention as described above. If the results are positive, 52, meaning the Al model described herein identified a gene expression pattern that is consistent with cancer, then the patient will be correctly stratified for further testing through established protocols including CT, XR, or MRI, as shown at 56. If the results are negative, 54, meaning the Al model described herein did not identify a gene expression pattern that is consistent withcancer, then the patient will be stratified into an annual screening pool, 58, for subsequent repeat blood test at a later date (e.g. the following year)

[0108] The methods and systems described above may be used for monitoring the progression of cancer, which is also illustrated in FIG. 3. As shown, patients who are screened for positive test results, as shown at 60, following the further testing 56 and / or are otherwise diagnosed with cancer may undergo routine blood collection and analysis steps, 62, at any suitable time points or intervals (for example, every 1, 4, 6, or 12 months). The sample collections times and intervals may be varied in various ways to meet the specific needs.

[0109] The blood samples are processed and prepared for RNA sequencing as described above. Specifically, as above, a preprocessor module is run to demultiplex and process FASTQ files into read counts and the DiffEx™ module is run to manipulate the read counts and calculate Iog2 fold change values to determine gene expression for all samples. The Iog2 fold change values are used as input in the stacking ensemble Al model described above, which determines if one or more of the samples exhibit a gene expression pattern that is indicative of cancer. A report is generated for each time point which will indicate if a cancer signal was detected in the collected time point which will convey disease progression, response to therapy, or other valuable insights.

[0110] Although the above description includes reference to certain specific embodiments, various modifications are possible. Any examples provided herein are included solely for the purpose of illustration and are not intended to be limiting in any way. Any drawings provided herein are solely for the purpose of illustrating various aspects of the description and are not intended to be drawn to scale or to be limiting in any way. The scope of the embodiments described herein should not be limited by the preferred embodiments set forth in the above description but should be given the broadest interpretation consistent with the present specification as a whole. The disclosures of all references in the present description herein are incorporated herein by reference in their entirety.

[0111] Various terms are used in the present description. Where appropriate, such terms are assigned meanings as indicated below or as they are introduced for the purpose describing one or more embodiments.

[0112] The term “embodiment” is used herein to describe one or more examples of representations or implementations of one or more features, elements, structures, or characteristics etc. (collectively “features”) of the present description. It will be understood that the features of a given embodiment of the description are not necessarily limited to such embodiment. In other words, any of the features described herein with respect to one embodiment may be used with, incorporated into, or combined with any of the described embodiments.

[0113] The terms “early-stage” and “late-stage” are used herein with respect to stages of a cancer state. As used herein, these terms are ascribed their known meanings. “Early-stage” refers to stage l-ll, and “late-stage” refers to stage lll-IV.

[0114] As used herein, the terms “classifier”, “classifier model”, or “classification model” refer to computer implemented machine learning algorithms and may be used interchangeably. Examples of classifiers include support vector machine(s) (“SVM”), AdaBoost classifier(s), penalized logistic regression, elastic nets, regression tree system(s), gradient tree boosting system(s), naive Bayes classifier(s), neural nets, Bayesian neural nets, k-nearest neighbor classifier(s), and random forests. As discussed further below, in a preferred embodiment, the description contemplates the use of one or more specific classifiers. Further descriptions of preferred classifiers are provided herein.

[0115] A “classification system” will be understood to mean a machine learning system that executes at least one classifier. Classifiers are “trained” using machine learning systems by building a model from inputs, finding relationships between variables (using mathematical techniques), and generating a prediction as an output.

[0116] As used herein “machine learning” refers to algorithms that comprise a series of executable statements encoded on a machine readable medium. When a machine learning algorithm is executed by a computer, the computer has the ability to learn without being explicitly programmed. Such algorithms include those that learn from and make predictions about data. Various machine learning algorithms are known in the art and a number are described further herein. Some examples include, but are not limited to, decision tree learning, neural network, deep learning neural network, support vector machines, rule base machine learning, random forest, logistic regression, pattern recognition algorithms, etc.

[0117] The term “computer” will includes at least one hardware processor that uses, or has access to, at least one memory.

[0118] As used herein, the term “algorithm” will be understood to mean series of encoded computer-readable instructions stored, either permanently or temporarily, on at least one memory that is accessible by at least one processor of a computer. The at least one processor executes the instructions that are stored in the at least one memory to process data.

[0119] As used herein, the term “neural network” will be understood to mean a computational model inspired by the human brain's network of neurons. It consists of layers of interconnected nodes (neurons) that process input data to produce an output. Neural networks are particularly good at capturing non-linear relationships in data.

[0120] The term “Rectified Linear Unit” (ReLU) will means an activation function defined by the equation f(x) = max(0, x). If the ReLU function received a negative input, the function will output zero, but will report the output directly if the input is positive. ReLU introduces nonlinearity into the model, enabling it to learn complex patterns. ReLU is computationally efficientand helps prevent issues like the vanishing gradient problem, where gradients become too small for effective learning in deep networks.

[0121] The “sigmoid function” is defined by the equation:1ff(%) = , , _1 + exx

[0122] The sigmoid function maps any input value to an output value in the range of 0 to 1. In the context of the embodiments described herein, the output value of the sigmoid function, when applied, represents the probability of lung cancer presence. In particular, an output close to 1 indicates a high probability of lung cancer, while an output close to 0 indicates a low probability.

[0123] A “Binary Cross-Entropy Loss” function refers to a function used in binary classification tasks that measures the difference between the predicted probabilities and the actual labels. It is defined as:Loss= -j^ytx los(p yD) + (1 - yd x iog(i - p(yD)i = l

[0124] where: y; is the true label (e.g., 0 for no cancer, 1 for cancer); p(yi) is the predicted probability from the model; and N is the number of samples. This function is suitable for binary classification tasks where the outcome is either 0 or 1 (e.g., in the context of the present description, the absence (0) or presence (1) of lung cancer). The binary cross-entropy loss function measures the difference between the predicted probabilities and the actual labels. It penalizes confident but wrong predictions more than less confident ones, thus guiding the model to improved accuracy.

[0125] The term “Adam”, or Adaptive Moment Estimation, will be understood to be an adaptive learning rate optimization algorithm that combines the advantages of two other methods, AdaGrad and RMSProp, as known in the art. Adam adjusts the learning rate during training for each parameter, which leads to faster convergence and better performance. Adam adjusts the learning rate individually for each parameter based on estimates of first and second moments of the gradients, leading to faster convergence.

[0126] The term “Dropout Layer” will be understood to mean a regularization technique where a fraction of input units is randomly set to zero during training. This prevents units from co-adapting too much and reduces overfitting. By training different subsets of the network on different data, dropout can improve the network's ability to generalize to new data.

[0127] The terms “detection” and “diagnosis” as used herein will be understood to have the same meaning. Thus, detection of a disease state by the methods described herein will be considered as a diagnosis of such disease.

[0128] The terms “comprise”, “comprises”, “comprised” or “comprising” may be used in the present description. As used herein (including the specification and / or the claims), and unlessstated otherwise, these terms are to be interpreted as open-ended terms and as specifying the presence of the stated features, integers, steps or components, but not as precluding the presence of one or more other feature, integer, step, component or a group thereof as would be apparent to persons having ordinary skill in the relevant art. Thus, the term "comprising" as used in this specification means "consisting at least in part of”. When interpreting statements in this specification that include that term, the features, prefaced by that term in each statement, all need to be present but other features can also be present. Related terms such as "comprise" and "comprised" are to be interpreted in the same manner.

[0129] The phrase “consisting essentially of” or “consists essentially of” will be understood as generally closed terms, with the exception of allowing inclusion of additional items, materials, components, steps, or elements, that do not materially affect the basic and novel characteristics or function of the item(s) used in connection therewith. For example, trace elements present in a composition, but not affecting the composition's nature or characteristics would be permissible if present under the “consisting essentially of” language, even though not expressly recited in a list of items following such terminology. When using an open-ended term, such as “comprising” or “including”, it will be understood that direct support should be afforded also to “consisting essentially of’ language as well as “consisting of” language as if stated explicitly and vice versa. In essence, use of one of these terms in the specification provides support for all of the others.

[0130] For the purposes of the present description, and unless otherwise indicated, all numbers expressing quantities, percentages or proportions, and other numerical values used in the specification and claims, are to be understood as being modified in all instances by the term “about.” Accordingly, unless indicated to the contrary, the numerical parameters set forth herein are approximations that may vary depending upon the desired properties sought to be obtained by the present invention, inclusive of the stated value and has the meaning including the degree of error associated with measurement of the particular quantity. The term “about” generally refers to a range of numbers that one of ordinary skill in the art would consider as a reasonable amount of deviation to the recited numeric values (i.e. , having the equivalent function or result). For example, the term “about” can be construed as including a deviation of ±10 percent of the given numeric value provided such a deviation does not alter the end function or result of the value. Therefore, a value of about 1% can be construed to be a range from 0.9% to 1.1%.

[0131] The term "and / or" can mean "and" or "or". Unless stated otherwise herein, the articles “a” and “the”, when used to identify an element, are not intended to constitute a limitation of just one and will, instead, be understood to mean “at least one” or “one or more”.

[0132] From a practical perspective, a physical implemented product (more like a service) is the lung cancer testing service. Another possible variation can include a physical,- at-home testing kit. This kit can be coupled with a mobile blood collection service that includes sending a phlebotomist to the patient’s home. In another variation, there can be a kit where a patient cancollect their own blood using a finger prick device or a micro-sampling device. In another variation, a clinical lab testing kit can be provided. The previously described kit that includes everything for a clinic or hospital to collect a patient’s blood for testing.

[0133] This apparatus can be sold to clinics and hospitals that will perform all aspects of blood collection, sample processing, and sequencing. The kit can include disposable preloaded cartridges for the device. These are single-use cartridges that are designed for one test and contain all the buffers and reagents required. There can be different types of cartridges for different types of testing. A user may obtain access to the OncoSage online portal (SaaS model). Cloud-hosted OncoSage portal that provides access to the Al model and results dashboard. The portal can include OncoSage risk reports. Add-on upgrade for a patient to receive personalized risk reports based on their testing results. In another variation, there can be a monitoring subscription. For current cancer patients or cancer survivors so they can purchase vouchers or a subscription that will allow for testing every 6 months or every 12 months with complete results.

[0134] The claimed invention addresses a real-world biomedical problem - early-stage cancer detection - through a concrete, physical transformation of a biological sample into a clinical decision. Unlike conventional diagnostic algorithms that merely interpret pre-existing data, the OncoSage platform initiates with a biological sample input (specifically, a patient’s blood sample) and transforms it through physically verifiable laboratory steps, including platelet isolation, RNA extraction, and sequencing, culminating in a real-time clinical classification output. The invention produces a discernible physical effect by transforming raw biological material into a clinically actionable diagnostic classification, thereby influencing downstream medical decision-making. Therefore, the OncoSage Al-driven classification described herein does not merely process abstract data; it operates in coordination with physical systems, such as but not limited to sequencing instruments, laboratory automation platforms, and computing infrastructure, to generate a materially significant change: the detection of cancer presence, type, and stage from patient-derived platelet RNA obtained from a blood sample.

[0135] The actual invention resides not merely in the algorithmic model, but in the synergistic interaction between physical blood sample processing, sequencing instrumentation, and the Al models, resulting in a materially transformed diagnostic capability. Overall, the system as claimed provides a technical improvement in biomarker detection, reducing the required sequencing depth, memory usage, and processing time by optimizing transcriptomic signal extraction from low-input platelet RNA samples.

[0136] Exemplary Embodiments.

[0137] Lung cancer screening (see Fig 2). A patient undergoes a routine blood collection to collect a blood sample. The blood sample is sent to a processing lab. Platelets are isolated from the blood sample using standard procedures which would be known to someone familiar to theart. RNA is extracted from platelets and a library is prepared for RNA-sequencing. RNA-sequencing is performed to obtain FASTQ files for all samples. Preprocessor module is run to process FASTQ files into read counts. DiffEx module is run to manipulate the read counts and calculate Iog2 fold change values to determine gene expression. Log2 fold change values are used as input in the stacking ensemble Al model. Stacking ensemble Al model determines if the sample exhibits a gene expression pattern that is indicative of lung cancer. A report is generated for the sample to describe the results to the ordering clinician and / or patient.

[0138] Population-wide cancer screening (see Fig. 2). Multiple patients undergo routine blood collections to collect blood samples. The blood samples are sent to a processing lab. Platelets are isolated from blood samples using standard procedures which would be known to someone familiar to the art. RNA is extracted from platelets and a library is prepared for RNA-sequencing. Library for each sample is barcoded with a unique sequence for each patient sample. Multiplexed RNA-sequencing is performed to obtain FASTQ files for all samples.Preprocessor module is run to demultiplex and process FASTQ files into read counts DiffEx module is run to manipulate the read counts and calculate Iog2 fold change values to determine gene expression for all samples. Log2 fold change values are used as input in the stacking ensemble Al model. Stacking ensemble Al model determines if the samples exhibit a gene expression pattern that is indicative of cancer. Reports are generated for each patient sample to describe the results to the ordering clinicians and / or patients.

[0139] Screening at-risk populations (see Fig. 3). Patients who meet certain criteria (e.g. at-risk individuals) for cancer testing undergo routine blood collection to collect blood samples. Blood samples are processed and evaluated on the invention as described above. If the results are positive, meaning the invention’s Al models identified a gene expression pattern that is consistent with cancer, then the patient will be correctly stratified for further testing through established protocols including CT, XR, or MRI. If the results are negative, meaning the invention’s Al models did not identify a gene expression pattern that is consistent with cancer, then the patient will be stratified into an annual screening pool for subsequent repeat blood test at a later date (e.g. the following year).

[0140] Continuous monitoring of cancer progression (see Fig. 3). Patients who are screened for positive test results with the invention and / or are diagnosed with cancer may undergo routine blood collection to collect blood samples at time points (for example, every 1 , 4, 6, or 12 months). The blood samples are sent to a processing lab and processed and prepared for RNA sequencing as described above. Preprocessor module is run to demultiplex and process FASTQ files into read counts. DiffEx module is run to manipulate the read counts and calculate Iog2 fold change values to determine gene expression for all samples. Log2 fold change values are used as input in the stacking ensemble Al model. Stacking ensemble Al model determines if the samples exhibit a gene expression pattern that is indicative of cancer. A report is generatedfor each time point which will indicate if a cancer signal was detected in the collected time point which will convey disease progression, response to therapy, or other valuable insights.

[0141] Stacking Ensemble with Multi-Cancer Training Strategy - Lung Cancer Detection Example

[0142] Detecting lung-specific cancer signals from blood presents substantial technical challenges that span multiple domains of computational biology and machine learning. The high dimensionality of gene expression data creates significant computational and statistical challenges. RNA sequencing from blood samples generates measurements for tens of thousands of genes, yet sample sizes typically number in the thousands at most. This unfavorable ratio of features to samples creates the well-known "curse of dimensionality" where traditional machine learning approaches become unreliable due to the increased risk of overfitting and the exponential growth in the volume of the feature space.

[0143] Class imbalance naturally occurs in screening populations where lung cancer cases comprise a minority of the tested population, even in high-risk cohorts. This imbalance can bias classifiers toward predicting the majority class and requires careful consideration in model training, evaluation metrics, and threshold selection. Biological noise is inherent in gene expression measurements due to both technical and biological factors. Technical variation arises from sample collection protocols, RNA extraction methods, library preparation, and sequencing platforms. Biological variation stems from individual patient heterogeneity, circulating RNA degradation rates, and the stochastic nature of transcription itself.

[0144] The cancer specificity challenge represents a critical distinction from traditional binary classification problems. Unlike simpler approaches that distinguish "cancer" from "healthy," a clinically deployable lung cancer screening test must distinguish lung cancer specifically from both healthy individuals and patients with other malignancies. Generic cancer biomarkers such as proliferation markers or general inflammation signals may elevate in many cancer types and would lead to false positives if used to screen specifically for lung cancer. Clinical requirements demand extremely high sensitivity to avoid missing true lung cancer cases, as false negatives in cancer screening can have fatal consequences. Simultaneously, the test must maintain specificity to lung cancer to prevent inappropriate diagnostic workups and avoid flagging patients with other cancer types for lung-specific follow-up procedures.

[0145] The Multi-Cancer Discrimination Problem - A model was developed to address three-way discrimination problems that extends beyond traditional binary classification. The model must first detect lung cancer with high sensitivity to minimize false negatives in a screening context where missing a cancer diagnosis can result in disease progression and increased mortality. Second, it must avoid flagging healthy individuals to maintain acceptable specificity and reduce the burden of unnecessary follow-up procedures including CT scans and biopsies that carry their own risks and costs. Third, and most critically for clinical deployment, it must notflag other cancer types as lung cancer, thereby achieving cancer-type specificity rather than merely detecting general malignancy. This three-way distinction between lung cancer, healthy individuals, and other cancer types represents a significantly more challenging classification problem than standard binary discrimination. However, it is clinically essential for a lung cancerspecific screening test that can be deployed in real-world populations where undiagnosed nonlung malignancies may be present. A test that flags glioma or melanoma patients as having lung cancer would lead to inappropriate clinical workups, patient anxiety, and healthcare resource misallocation. Rationale for Ensemble Methods. Model-specific biases arise from the inductive biases built into each algorithm's learning procedure. By combining models with different biases, the ensemble can achieve more complete coverage of the pattern space. Ensemble decisions demonstrate improved generalization compared to single-model predictions because the averaging or meta-learning process reduces variance. Random errors made by individual models tend to cancel out in the ensemble, while systematic patterns that multiple models agree upon are reinforced. This variance reduction is particularly valuable in high-dimensional genomic data where individual model predictions may be unstable due to the many correlated features and limited samples. For clinical decision-making, consensus predictions from multiple models are inherently more trustworthy than single-model outputs. When multiple diverse algorithms independently arrive at the same conclusion, confidence in that conclusion is substantially higher than when relying on a single model that may have idiosyncratic failure modes. This increased reliability is essential for clinical deployment where prediction errors can have serious consequences for patient outcomes.

[0146] Dataset Characteristics and Preprocessing - Training Dataset Composition. Training cohort comprises 3,265 blood samples with paired RNA sequencing data and clinical annotations. The dataset includes 1,070 lung cancer cases representing 32.8% of the total samples. Critically, 695 samples from patients were included with other cancer types, representing 21.3% of the cohort. These other cancers include 414 glioma cases, 101 head and neck cancer cases, 92 esophageal cancer cases, 53 melanoma cases, and 35 prostate cancer cases. Healthy control samples number 1,500, representing 45.9% of the cohort. Gene expression measurements encompass 57,000 transcripts quantified through RNA sequencing of whole blood samples. Training Paradigm. Unlike traditional binary classification approaches that train on lung cancer versus healthy controls, a three-class training paradigm was implemented collapsed into binary output. The positive class consists of lung cancer samples only. The negative class comprises both healthy controls and other cancer types. This design choice represents a significant methodological innovation with important implications for model behavior and clinical utility. Traditional lung cancer detection models trained exclusively on lung cancer versus healthy controls learn to recognize generic cancer biomarkers that differentiate diseased from non-diseased states. Such models detect cancer-associated gene expressionchanges including upregulation of proliferation markers, cell cycle genes, and general stress responses. However, these signals are not specific to lung cancer and would be elevated in many malignancy types. When deployed in clinical screening populations that may include individuals with undiagnosed non-lung cancers, traditional models would generate false positives by flagging these other malignancies as lung cancer.

[0147] The three-class approach explicitly addresses this limitation by including other cancer types in the negative class during training. The model must learn to distinguish lungspecific transcriptom ic signatures from both healthy baselines and cancer signals originating from non-lung tissues. This fundamentally changes what the model learns. Instead of simply detecting "cancer versus no cancer," it must identify patterns unique to lung malignancies while ignoring or down-weighting generic cancer biomarkers that appear across multiple cancer types. The training dynamics reflect this increased complexity. In initial training, the model achieved approximately 80% accuracy by learning to separate lung cancer from healthy controls using general cancer markers. However, this strategy fails for the other cancer samples in the negative class, which also express these generic markers. To improve beyond 80% accuracy, the model must refine its learned features to focus on lung-specific patterns. This refinement process during later training epochs drives the final performance gains and produces a classifier with true cancer-type specificity. From a biological perspective, the model learns to focus on lung tissue-specific gene expression patterns including lung-specific transcription factors and differentiation markers. It identifies tumor microenvironment signatures unique to lung tissue including interactions between tumor cells and lung-resident immune populations and stromal cells. The model captures immune response patterns specific to lung malignancies, which differ from immune responses to brain tumors or melanoma due to tissue-specific antigen presentation and local cytokine milieu. Additionally, it detects circulating tumor-derived RNA species that originated specifically from lung tissue based on tissue-specific transcript variants and expression profiles. Rather than relying on generic cancer markers, the trained model down-weights signals such as general proliferation markers including MKI67 and PCNA that are elevated in all rapidly dividing malignancies. It ignores non-specific inflammation signals and systemic stress responses. It filters out apoptosis markers common across many pathological conditions. This specificity represents the key innovation enabling clinical deployment.

[0148] Clinical Validation of the Multi-Cancer Approach. The model achieved 90.9% specificity on this negative class that includes both healthy individuals and non-lung cancer patients. This demonstrates that the model successfully learned lung-specific signatures rather than generic cancer biomarkers. If trained only against healthy controls, one would expect substantially lower specificity when tested on a negative class including other malignancies, as the model would flag these other cancers as positive predictions.

[0149] In actual screening populations, individuals may harbor undiagnosed cancers of nonlung origin. A traditional model would flag these individuals as having lung cancer, leading to inappropriate lung-specific diagnostic workups including thoracic CT scans and potentially bronchoscopy or lung biopsies. The three-class model correctly classifies these individuals as not having lung cancer, allowing their actual conditions to be identified through appropriate diagnostic pathways. This cancer-type specificity represents the critical difference between a research demonstration and a clinically viable screening test.

[0150] Feature Selection from High-Dimensional Genomic Data - With 57,000 gene expression features but only 3,265 samples, this present the challenge of dimensionality where the number of features dramatically exceeds the number of observations by more than seventeen-fold. This unfavorable ratio leads to several well-documented pathologies that compromise both model performance and biological interpretability. Overfitting becomes nearly inevitable as models have sufficient capacity to memorize training data noise rather than learning generalizable patterns that reflect true biological signals. The computational burden grows substantially as the feature space expands, with many algorithms exhibiting polynomial or exponential scaling in the number of features, making training prohibitively slow or memoryintensive. Noise amplification occurs because the vast majority of genes provide no discriminatory information for lung cancer detection and merely add random variation that obscures true signals. Poor interpretability emerges as understanding which biological processes and pathways drive predictions becomes impossible when thousands of weakly-informative features contribute marginally to model decisions. A solution was developed employing a multi-stage feature selection pipeline that combines complementary methodologies, each addressing different aspects of the feature selection problem. This ensemble approach to feature selection parallels the ensemble approach to classification, leveraging the strengths of multiple methods while mitigating their individual weaknesses.

[0151] Stage One: Variance-Based Filtering - To begin, variance filtering to remove genes showed minimal expression variation across samples. Genes with near-constant expression provide no discriminatory information as they cannot distinguish between lung cancer and negative class samples. These invariant genes typically represent either constitutively expressed housekeeping genes required for basic cellular functions, genes expressed at undetectable levels in blood samples regardless of disease status, or technical artifacts from sequencing noise and background signal. The variance of each gene's expression was computed across all training samples and remove genes falling below the 25th percentile of the variance distribution. This conservative threshold eliminates approximately 14,250 genes (25% of 57,000) that are effectively uninformative while retaining all genes with moderate to high variance that might contribute to classification. The variance-based approach is computationallyefficient, requires no distributional assumptions, and serves as an effective first-pass filter that reduces the feature space for subsequent more computationally intensive selection methods.

[0152] Stage Two: Statistical Significance Testing for Differential Expression - Following variance filtering, statistical significance testing was applied to identify genes with differential expression between lung cancer cases and the combined negative class comprising healthy controls and other cancer types. This analysis addresses the question of which genes show systematic expression differences associated with lung cancer status rather than random fluctuations. Non-parametric tests was employed appropriate for RNA-sequencing count data that do not assume normal distributions. Specifically, the Wilcoxon rank-sum test was used to compare expression distributions between lung cancer and negative class samples for each gene. This test evaluates whether the distribution of expression values differs systematically between groups while making minimal distributional assumptions and maintaining robustness to outliers common in genomic data.

[0153] To address multiple testing concerns arising from testing thousands of hypotheses simultaneously, the Benjamini-Hochberg false discovery rate correction was applied to control the expected proportion of false positives among genes called significant. The false discovery rate threshold was set at 0.05, accepting that up to 5% of selected genes may represent false positives while ensuring that the vast majority reflect true differential expression. This procedure identifies approximately 8,500 genes showing statistically significant differential expression after multiple testing correction. These genes represent candidates whose expression patterns systematically differ between lung cancer and the negative class, though statistical significance alone does not guarantee predictive utility for classification.

[0154] Stage Three: Mutual Information Analysis - Mutual information provides a complementary perspective by quantifying the predictive power of each gene for the classification target. Unlike correlation-based methods that capture only linear relationships, mutual information measures the reduction in uncertainty about cancer status given knowledge of a gene's expression level, thereby capturing both linear and non-linear associations. For a gene X and cancer status Y, mutual information is defined as the Kullback-Leibler divergence between the joint distribution P(X,Y) and the product of marginals P(X)P(Y). High mutual information indicates that knowing the gene's expression provides substantial information about cancer status, while mutual information near zero indicates independence where the gene provides no predictive value. Continuous expression values were discretized into bins to enable mutual information estimation, using an adaptive binning strategy that preserves information content while maintaining sufficient samples per bin for reliable estimation. For each gene, mutual information was computed with the binary cancer status outcome and rank genes by their mutual information scores. Genes exceeding a threshold corresponding to the 90th percentile of the mutual information distribution were retained, selecting approximately 5,700genes with the highest predictive power. This approach captures genes with strong univariate predictive ability regardless of whether their relationship with cancer status is linear or exhibits more complex non-linear patterns such as threshold effects or non-monotonic relationships.

[0155] Stage Four: Correlation-Based Redundancy Reduction - The previous stages select genes individually based on their marginal properties without considering relationships among genes. Many genes exhibit high correlation due to co-expression in the same biological pathways, regulation by common transcription factors, or membership in tightly coupled gene regulatory networks. Including many highly correlated genes provides redundant information that does not improve model performance while increasing computational cost and reducing interpretability. Correlation-based redundancy reduction were implemented to identify and remove such redundant genes while preserving the most informative representative from each correlated cluster. The Pearson correlation coefficient were computed between all pairs of genes surviving the previous filtering stages, resulting in a correlation matrix of dimensions approximately 5,700 by 5,700. Gene pairs with absolute correlation exceeding 0.85 were identified, indicating near-perfect positive or negative linear relationships. For each correlated pair, the gene with higher mutual information with the cancer status outcome was retained and the gene with lower predictive power was removed. This greedy approach iteratively processes correlated pairs in order of decreasing correlation magnitude until no pairs exceed the threshold. The procedure removes approximately 1 ,200 highly redundant genes while retaining approximately 4,500 genes that provide relatively independent information. This reduces model complexity and improves interpretability by ensuring that selected genes represent diverse biological signals rather than multiple measurements of essentially the same underlying process.

[0156] Stage Five: Stability Selection Across Data Subsamples - Stability selection addresses the concern that feature selection on a single realization of the data may be unstable, with small perturbations in the sample leading to substantial changes in the selected gene set. Genes selected consistently across multiple random subsamples of the data are more likely to reflect robust biological signals rather than spurious associations that happen to appear strong in the particular training set. Stability selection was implemented by generating 100 bootstrap samples, each created by randomly sampling 80% of the training data with replacement while maintaining class proportions. For each bootstrap sample, the entire feature selection pipeline was performed from variance filtering through redundancy reduction. This produces 100 potentially different selected gene sets, each representing the features that would be selected if that bootstrap sample had been the actual training data. A stability score was computed for each gene as the proportion of bootstrap iterations in which that gene was selected. Genes with stability scores near 1.0 are selected consistently across nearly all data realizations, indicating robust association with cancer status that does not depend sensitively on which particularsamples are included. Genes with low stability scores are selected only in a minority of bootstrap samples, suggesting their apparent importance may result from overfitting to specific sample compositions rather than genuine biological signal. Genes with stability scores exceeding 0.70 were retained, meaning they are selected in at least 70 of the 100 bootstrap iterations. This stability threshold balances retaining genes with robust evidence of importance while maintaining sufficient features for powerful classification. Approximately 1,800 genes exceed the stability threshold, representing the intersection of statistical significance, predictive power, low redundancy, and consistent selection across data perturbations.

[0157] Stage Six: Model-Based Feature Importance Ranking - This stage employs modelbased feature selection using gradient-boosted decision trees to directly evaluate each gene's contribution to predictive performance in the context of all other genes. Unlike the previous univariate methods that evaluate genes individually or in pairs, this approach considers multivariate feature importance accounting for interactions and conditional dependencies. This was implemented using the SelectFromModel meta-transformer from scikit-learn's feature selection module with XGBoost as the base estimator. XGBoost provides feature importance scores based on multiple criteria including gain measuring the average improvement in prediction accuracy when splitting on each feature, cover measuring the average number of samples affected by splits on each feature, and frequency counting how many times each feature is used for splitting across all trees. The SelectFromModel procedure operates through three steps. First, XGBoost classifier was trained with 300 trees, maximum depth 6, and learning rate 0.05 on the training data using the 1,800 genes surviving stability selection. This training process implicitly ranks genes by their contribution to predictive performance through the tree-building algorithm's greedy feature selection at each split point. Genes that consistently improve node purity and prediction accuracy receive high importance scores, while genes that are never selected for splitting or provide only marginal improvements receive low scores. Second, the distribution of feature importance scores across the 1,800 genes was examined and established a threshold for selection. Rather than using a fixed absolute threshold that might be too lenient or too stringent depending on the importance distribution, a relative threshold was used based on the mean importance score. Specifically, genes were selected with importance scores exceeding 1 ,5X the mean importance, corresponding approximately to the top 33% of features. Third, the feature space was transformed by retaining only the genes deemed important by the XGBoost importance ranking, which produces the final selected feature set. This model-based selection identifies 505 genes as the most significant features for distinguishing lung cancer patients from the negative class comprising healthy controls and other cancer types. These genes represent those that XGBoost, through its tree-building procedure, determined to be most informative for classification when considering feature interactions and conditional relationships. The importance ranking reflects not just marginal association with cancer status but contributionto predictive accuracy in the context of other features. Genes that might show moderate univariate associations can receive high importance if they provide unique information not captured by other genes, while genes with strong marginal associations may receive lower importance if their information is redundant with other selected features.

[0158] Biological and Statistical Properties of Selected Features - The reduction from 57,000 initial genes to 505 final features represents a 94% dimensionality reduction that provides multiple benefits. The resulting feature-to-sample ratio of 505 features to 2,775 training samples improves to approximately 1:4.6, moving closer to the rule-of-thumb guideline of at least 5-10 samples per feature for reliable model training. This improved ratio substantially reduces overfitting risk compared to the original 1:0.06 ratio where features outnumbered samples. Computational efficiency improves dramatically as training time scales roughly with the square of feature count for many algorithms, so the 94% reduction translates to approximately 97% reduction in computational cost for quadratic algorithms and even greater savings for higher-order algorithms.

[0159] Stacking Ensemble Methodology - Prevention of Data Leakage Through Out-of-Fold Predictions. A critical challenge in stacking implementation is avoiding data leakage that would allow the meta-model to overfit. If base models were trained on the full training set and then train the meta-model on predictions from those same samples, the meta-model has access to information it should not possess. The base models have already seen these samples during training, so their predictions are overly optimistic and do not reflect true generalization performance. A meta-model trained on such predictions would learn to exploit overfitting artifacts in the base models rather than learning genuine combination strategies.

[0160] A solution was presented that employs out-of-fold predictions generated through cross-validation. The training was partitioned set into five stratified folds maintaining class proportions. For each fold, all base models were trained on the remaining four folds comprising 80% of the training data. Then predictions were generated for the held-out fold using these models that have never seen these particular samples. These predictions represent out-of-fold predictions where the base models are predicting on samples they did not train on, accurately reflecting their generalization behavior. By repeating this process for all five folds, out-of-fold predictions were obtained for every sample in the training set. These predictions serve as metafeatures fortraining the meta-model. Critically, each sample's meta-features come from base models that never saw that sample during their training, ensuring no information leakage. The meta-model learns to combine base model predictions as they would appear on truly unseen data rather than overfitted training data. After meta-model training on out-of-fold predictions, all base models were trained on the complete training set to leverage all available data for maximum performance. These fully-trained base models generate predictions on the test set, which the meta-model then combines to produce final predictions. This two-stage processensures both that the meta-model trains on realistic base model behaviors and that final predictions use base models trained on maximum data. Six base classifiers were selected representing diverse algorithm families to maximize ensemble performance through complementary error patterns and diverse inductive biases: AdaBoost (Adaptive Boosting), XGBoost (Extreme Gradient Boosting), LightGBM (Light Gradient Boosting Machine), Random Forest, Support Vector Machine, and Custom Neural Network.

[0161] Custom Neural Network Meta-Model Architecture. Unlike standard stacking implementations that employ logistic regression as the meta-learner, a custom neural network architecture was designed specifically optimized for learning how to combine base model predictions. This represents a significant departure from conventional practice and provides substantial performance advantages.

[0162] Logistic regression, while interpretable, implements only linear combinations of base model predictions. It can learn fixed weights for each model but cannot capture non-linear interaction effects such as conditional trust relationships where "if model A predicts X and model B predicts Y, then the final prediction should be Z." In genomic applications where base models may be confident in different regions of feature space, these conditional relationships are critical. The developed meta-model neural network comprises four layers implementing gradual dimensionality reduction from six base model predictions to a single final prediction. The first layer contains 32 neurons accepting the six base model probability outputs as input. This relatively large first layer compared to the six inputs allows learning of numerous pairwise and higher-order interactions between base model predictions. The second layer reduces to 16 neurons, learning compressed representations of the first-layer interaction features. The third layer further reduces to 8 neurons, capturing the most salient patterns for final prediction. The output layer contains a single neuron with sigmoid activation producing the final probability.

[0163] LeakyReLU activations with negative slope 0.1 enable gradient flow even for negative inputs, preventing the "dead neuron" problem where neurons get stuck with zero activation and stop learning. Batch normalization after the first two layers stabilizes training by normalizing layer inputs, reduces sensitivity to learning rate selection, and acts as mild regularization. Dropout regularization with decreasing rates (0.3, 0.2, 0.2 across layers) prevents overfitting by randomly deactivating neurons during training, forcing the network to learn robust patterns that do not depend on any single neural pathway. L2 weight regularization with coefficient 0.001 penalizes large weights, promoting simpler models and reducing overfitting. This architecture supports context-dependent weighting where the meta-model learns different model combination strategies for different input regions. For example, it might learn that when tree-based models all predict cancer with high probability, this is a strong signal. However, when tree models are uncertain (probabilities near 0.5) but the SVM is confident, theSVM should be trusted more heavily. These conditional relationships emerge from the nonlinear transformations across layers.

[0164] The developed meta-model learns complex conditional logic such as "when XGBoost and LightGBM agree with high confidence, trust their prediction strongly," "when only SVM predicts cancer while tree models predict control, investigate the neural network prediction as a tiebreaker," "when all models disagree substantially, flag the sample as uncertain and default to the conservative side (predicting cancer to avoid false negatives)," and "when all models agree unanimously, be very confident in the prediction regardless of absolute probability values."

[0165] Training employs binary cross-entropy loss optimized via the Adam algorithm with learning rate 0.001. Up to 100 epochs was trained with early stopping monitoring validation loss to prevent overfitting. Learning rate reduction on plateau halves the learning rate when validation loss stops improving, allowing fine-grained optimization near local minima. Batch size of 32 provides a balance between gradient estimate stability and computational efficiency.

[0166] Compared to logistic regression with six base model inputs and one output requiring seven trainable parameters, the neural network employs approximately 700 parameters providing two orders of magnitude greater expressiveness. This capacity enables learning of subtle patterns in base model prediction combinations that simpler meta-learners miss.However, the extensive regularization through dropout, L2 penalties, batch normalization, and early stopping prevents this increased capacity from leading to overfitting.

[0167] Threshold Optimization - Traditional approaches to threshold selection suffer from several limitations. Using the default 0.5 threshold ignores clinical context and applicationspecific cost structures. Selecting thresholds based on maximizing overall accuracy implicitly weights classes by their prevalence, disadvantaging the minority cancer class. Choosing thresholds through manual inspection of curves lacks reproducibility and scientific justification. These limitations motivate present development of a scientifically rigorous, reproducible threshold selection methodology. Clinical Requirements and Constraints. For cancer screening applications, sensitivity or recall representing the proportion of true cancer cases correctly identified constitutes the paramount metric. Present clinical target specifies detecting at least 91% of cancer cases, as missing 9 or fewer cancers per 100 patients represents an acceptable trade-off given the screening context. However, it was also aimed to limit sensitivity to approximately 95% to maintain reasonable precision. While high sensitivity is critical, if achieving the last few percentage points of sensitivity requires a dramatic increase in false positives, the marginal benefit may not justify the cost in terms of unnecessary procedures and healthcare resource utilization. Precision or positive predictive value, representing the proportion of positive predictions that are true cancers, serves as a secondary but important consideration. Present target exceeds 80% precision, meaning at least 80% of individuals flagged by the test truly have lung cancer. This maintains a reasonable balance where mostpositive predictions are correct, limiting the burden of false positive follow-up while prioritizing the critical goal of detecting as many cancers as possible. The sensitivity-specificity trade-off reflects the fundamental reality that improving one metric typically degrades the other. Higher sensitivity catches more cancers but flags more healthy individuals. Presently developed approach explicitly quantifies this trade-off and makes principled decisions about where to operate on the trade-off curve based on clinical priorities that weight sensitivity more heavily than specificity. Comprehensive Threshold Evaluation. 100 threshold values were evaluated ranging from 0.20 to 0.70 in increments of 0.005, producing a fine-grained characterization of model performance across the relevant operating range. For each threshold, all standard classification metrics were computed including precision, recall, specificity, F1 score, and accuracy on the test set. This systematic sweep provides complete information about the sensitivity-specificity trade-off and reveals performance plateaus where small threshold changes produce minimal metric changes. Rather than standard F1 score that equally weights precision and recall, F-beta score was employed with beta parameter 1.2, which weights recall 1.2 times more heavily than precision. The F-beta formulation is (1 + betaA2) * (precision * recall) I (betaA2 * precision + recall), explicitly encoding the clinical priority for high recall. This metric aligns evaluation with clinical objectives, capturing the domain knowledge that in cancer screening, sensitivity outweighs specificity in importance. Candidate thresholds were identified achieving the target recall range of 91% to 95%. Among these candidates, F-1.2 scores were computed and the maximum achievable value was identified. Then a performance plateau was defined as all thresholds achieving F-1.2 scores within 1% of this maximum. This plateau concept acknowledges that small variations in F-beta within a narrow range likely reflect noise rather than meaningful performance differences and provides flexibility to select thresholds based on secondary criteria while maintaining near-optimal performance. Within the performance plateau, the threshold was selected. This value represents the approximate center of the typical optimal range for imbalanced classification problems and provides an interpretable reference point. If the plateau shifts due to data characteristics, this method automatically adapts by selecting the closest threshold within the optimal region. This rule-based selection within the plateau ensures reproducibility while acknowledging that exact threshold choice within a narrow optimal range is not critical. Present selected threshold of 0.4150 achieves 91.3% recall, 83.1% precision, 86.98% F1 score, 91.02% accuracy, and 90.88% specificity. This operating point successfully detects 91.3% of lung cancers while maintaining over 83% precision, meeting both the primary clinical requirement for high sensitivity and secondary goal of reasonable precision.

[0168] Advantages Over Alternative Approaches. Compared to using the default 0.5 threshold without clinical consideration, present method explicitly targets the clinically relevant recall range and achieves substantially higher sensitivity. The default threshold would yield approximately 87% recall, missing 13% of cancers compared to present 8.7% miss rate. Whilethe default threshold might achieve slightly higher precision, the trade-off is clearly unfavorable from a clinical screening perspective. This approach differs from simple metric maximization such as choosing the threshold that maximizes accuracy or F1 score without constraints. Pure accuracy maximization would likely produce a threshold favoring the majority class and achieving lower recall. By explicitly constraining the recall range and using F-beta scoring, domain knowledge about clinical priorities were incorporated directly into the optimization.

[0169] By identifying performance plateaus rather than single optimal points, this approach acknowledges measurement uncertainty and avoids overfitting threshold selection to test set noise. Small threshold variations within the plateau produce negligible performance differences, and selecting any threshold within the plateau would yield similar clinical outcomes. This robustness is important for ensuring that the exact threshold choice generalizes to new patient populations.

[0170] Results - Stacking Ensemble Performance at Optimal Threshold. At the selected threshold of 0.4150, the stacking ensemble achieves sensitivity of 91.3% detecting 146 of 160 lung cancer cases, precision of 83.1% where 146 of 176 positive predictions are true lung cancers, specificity of 90.9% correctly identifying 300 of 330 negative class samples, F1 score of 87.0% representing balanced harmonic mean of precision and recall, accuracy of 91.0% with 446 correct predictions of 490 total samples, and area under the ROC curve of 0.950 demonstrating excellent discrimination ability. The confusion matrix reveals 300 true negatives where negative class samples are correctly identified as not lung cancer, 30 false positives where negative class samples are incorrectly flagged as lung cancer, 14 false negatives where lung cancer cases are missed, and 146 true positives where lung cancer cases are correctly detected. The 30 false positives include both healthy individuals incorrectly flagged and other cancer types incorrectly identified as lung cancer. The false positive rate on the mixed negative class remains acceptable at 9.1%, and critically, the majority of other cancer samples are correctly not flagged as lung cancer, demonstrating cancer-type specificity.

[0171] From a clinical interpretation perspective, the model detects 91 of every 100 lung cancer patients missing only 9, which represents excellent performance for early-stage cancer screening. Of every 100 individuals flagged as having lung cancer, 83 truly have the disease and 17 are false alarms requiring follow-up to resolve. When the test indicates lung cancer, there is 83% probability of true disease. When the test indicates no lung cancer, there is 95.5% probability of truly not having lung cancer calculated as 300 / (300+14). The negative predictive value of 95.5% provides strong reassurance when the test is negative.

[0172] The Ensemble Advantage: Mechanisms of Improvement. Fig. 4 visualizes recall performance across models illustrates the ensemble advantage. AdaBoost achieves 70.8%, Random Forest 68.9%, SVM 72.7%, XGBoost 83.9%, LightGBM 83.9%, CustomNN 83.9%,while stacking achieves 91.3%, representing a 7.4 percentage point improvement over the best single model. This improvement stems from multiple complementary mechanisms.

[0173] Complementary error patterns mean different models misclassify different samples based on their algorithmic biases. A cancer sample that XGBoost misses due to locally noisy features might be correctly identified by SVM considering global patterns. A cancer sample that all tree-based models miss might be caught by the neural network that learned different feature representations. When combining diverse models, the probability that all models simultaneously misclassify the same sample is substantially lower than any individual model missing that sample. The meta-model learns to recognize these complementary patterns and trust the model that is likely correct for each specific case. Confidence calibration allows the meta-model to learn when to trust each base model. When multiple models agree with high confidence, this consensus provides strong signal. When models disagree, the meta-model has learned which model to weight more heavily based on characteristics of the disagreement pattern. For samples where all models are uncertain (probabilities near 0.5), the meta-model can recognize this uncertainty and adjust its combination strategy accordingly, perhaps defaulting to the conservative choice of predicting cancer to avoid false negatives. Noise reduction occurs because individual models may make random errors on specific samples due to idiosyncrasies of their learning procedure or initialization. By averaging information across models, random errors tend to cancel while systematic patterns that multiple models independently discover are reinforced. The ensemble acts as a form of model regularization, producing more stable predictions than any single model. Synergistic predictions emerge for samples at different points in the difficulty spectrum. Some cancer cases exhibit obvious genomic signatures where all models confidently predict cancer. For these cases, the meta-model reinforces the consensus with very high confidence. Some cancer cases show subtle signals where only some models detect cancer. For these cases, the meta-model learns specific combinations indicating cancer, such as "if SVM and neural network both predict cancer even though tree models are uncertain, this is likely a true cancer." Some control cases exhibit cancer-like features causing some models to misclassify. For these cases, the meta-model learns to recognize false alarm patterns and override spurious positive predictions from individual models.

[0174] Cancer-Type Specificity Validation. The present model achieves 90.9% specificity on a negative class that includes 85 other cancer samples along with 245 healthy controls. Of the 330 negative class samples, 300 are correctly identified as not lung cancer. This means the model correctly rejects approximately 82 of 85 other cancer samples (90.9% specificity applied to the estimated other cancer count). If trained only against healthy controls, the model would have learned generic cancer biomarkers that elevate in all malignancies. When tested on a negative class including other cancer types, such a model would likely flag most of these other cancers as positive, yielding substantially lower specificity potentially below 50% on othercancers. The observed 90.9% specificity on the mixed negative class demonstrates that the model successfully learned lung-specific cancer signatures rather than generic malignancy markers. In real screening populations, individuals may harbor undiagnosed gliomas, melanomas, or other cancers. A traditional model would inappropriately flag these individuals for lung cancer follow-up including thoracic CT imaging and potentially invasive procedures like bronchoscopy or lung biopsy. These procedures carry their own risks and costs while failing to address the patient's actual condition. The present model correctly identifies these individuals as not having lung cancer, allowing appropriate diagnostic workups to identify their actual malignancies. This cancer-type specificity represents the key innovation enabling clinical deployment in real-world heterogeneous populations.

[0175] Summary. The present inventors have developed a multi-cancer training paradigm where other cancer types were explicitly included in the negative class during model development. This represents a fundamental departure from traditional binary classification approaches that train on target cancer versus healthy controls. This approach has profound implications for what the model learns and whether it achieves cancer-type specificity versus generic cancer detection. A custom neural network architecture was specifically developed for meta-learning in stacking ensembles. A carefully regularized neural network substantially improves performance through its ability to learn non-linear combination rules. A hierarchical architecture with gradual dimensionality reduction enables learning of multi-level patterns. The first hidden layer with 32 neurons learns pairwise interactions between base models. The second layer with 16 neurons learns higher-order combinations of these interactions. The third layer with 8 neurons distills the most salient patterns. This progressive abstraction mirrors successful deep learning architectures in other domains. Extensive regularization through dropout, L2 penalties, and batch normalization prevents the increased model capacity from leading to overfitting. The out-of-fold training scheme ensures the meta-model trains on realistic base model behaviors rather than overfitted predictions. The combination of high capacity and strong regularization produces a meta-learner that captures complex patterns while generalizing well to unseen data.

[0176] Exemplary Device Configuration - Prior approaches for cancer detection based on transcriptomic analysis suffer from fundamental technical limitations when applied to blood-derived RNA. Cancer-associated transcriptional signals in peripheral blood are often weak, heterogeneous, and systemically distributed, and are further obscured by substantial biological variability and technical noise arising from sample handling, sequencing depth, and batch effects. Conventional normalization and analysis techniques, which rely on global expression assumptions or post-hoc batch correction, frequently suppress or distort clinically meaningful signal, resulting in unstable model performance and poor generalization across cohorts. The disclosed invention addresses these technical challenges by jointly engineering platelet-focusedbiological enrichment, physical sample-processing hardware, control-anchored molecular normalization, and machine-learning-based inference into a unified system optimized for robust cancer detection. Furthermore, prior approaches to blood-based cancer detection have primarily focused on circulating tumor DNA, methylation signatures, isolated biomarkers, or purely computational classifiers operating on globally normalized data. Such approaches do not physically enrich for tumor-influenced platelet populations, do not normalize molecular signals relative to contemporaneous control samples within the same processing context, and do not integrate physical sample-processing devices with machine-learning-based inference in a unified system. The disclosed invention departs from these approaches by combining platelet-focused biological enrichment, control-anchored normalization, and integrated hardwaresoftware operation, resulting in improved robustness, generalizability, and diagnostic sensitivity.

[0177] In some embodiments, a fully integrated diagnostic system is provided, combining one or more of biological sample processing, microfluidics, molecular stabilization, and computational inference all into one unified physical device. The system is configured to accept a small volume of physical platelet sample, such as whole blood sample, and produce structured molecular data representative of platelet-derived transcriptomic signals associated with cancer. The system may be deployed in clinical laboratories, outpatient clinics, pharmacies, or home environments. In one embodiment, the system is a single-use, disposable cartridge engineered to perform blood fractionation (i.e. using a centrifuge or a microfluidics chip), platelet enrichment (i.e. using a gradient chamber), molecular stabilization (i.e. using platelet and / or RNA stabilization buffers or agents), and data preparation (i.e. using RNA sequencing instruments or assay kits). In one embodiment, the cartridge comprises one or more of: 1) an input port configured to receive whole blood, 2) microfluidic channels configured to separate blood components by size, density, or flow characteristics, 3) gradient chambers that enrich platelet populations while excluding other cells (such as nucleated cells), 4) reservoirs containing proprietary buffers optimized for platelet preservation, 5) chambers preloaded with RNA-stabilizing reagents, and 6) sealed waste compartments. The cartridge is manufactured using biocompatible polymers and is hermetically sealed to prevent contamination or RNA degradation. In one embodiment, the cartridge is preloaded with lyophilized reagents. In some embodiments, the cartridge geometry is specifically designed to favor and preferentially retain tumor-educated platelets, which exhibit altered surface markers, activation stress, RNA cargo, or physical properties relative to baseline platelets. Upon introduction of blood into the cartridge, passive and / or active microfluidic mechanisms guide the sample through a series of separation steps. These may include but are not limited to: laminar flow separation, inertial microfluidics, deterministic lateral displacement, density-gradient channels, or antibody-coated regions that selectively bind platelet surface markers. The cartridge outputs a platelet-enriched fraction isolated within one or more chambers of the cartridge, without requiring external centrifugation.

[0178] Once platelets are isolated, the cartridge introduces stabilization buffers that immediately arrest RNA degradation and preserve transcriptomic integrity. In some embodiments, platelet lysis and RNA capture occur directly within the cartridge. In other embodiments, RNA remains encapsulated within intact platelets until downstream processing. In one embodiment, the cartridge comprise nucleic acid binding surfaces, capture beads, or chemical matrices configured to retain RNA molecules for subsequent amplification or sequencing. In some embodiments, the cartridges are cancer-specific. In one embodiment, the buffer composition of the cartridge is selected based on the target cancer to be diagnosed. In one embodiment, the microfluidic geometries of the cartridge is configured based on the target cancer to be diagnosed. In one embodiment, the cartridge comprise nucleic acid binding surfaces, capture beads, or chemical matrices configured to retain RNA molecules or biomarkers specific to the target cancer to be diagnosed. In one embodiments, the cartridge comprises specific stabilization buffers and / or protocols based on the target cancer to be diagnosed. In some embodiments, the cartridges are optimized based on the target cancer to be diagnosed. In some embodiments, the cartridge is inserted into a benchtop reader device, which provides mechanical, fluidic, and electronic control. This modular design allows different cancer-specific cartridges to be utilized based on the target cancer diagnosis in question, without requiring changes to the core reader device. The reader device may include but are not limited to: actuators that control fluid flow through the cartridge, sensors that monitor flow rate, pressure, temperature, or optical signals, onboard processors that record run metadata, and communication modules that transmit structured output data to the computer processors described herein (such as the OncoSage platform). The reader device does not need to perform sequencing itself. Instead, it produces a digitally encoded molecular profile that can be correlated with sequencing-based or reference- based transcriptomic measurements. The reader device generates machine-readable outputs representing the molecular state of the platelet-derived sample. In one embodiment, the machine-readable outputs contain structured data representing 1) quantitative representations of RNA abundance, 2) gene- or pathway-level signatures, 3) quality metrics, and / or 4) batch identifiers (such as time stamp). The machine-readable outputs are structured such that they can be directly processed by the processors described herein (such as OncoSage computational pipeline), including conducting normalization, feature transformation, and inference modules.

[0179] While certain embodiments are described with reference to RNA sequencing and gene-level read counts, the invention is not limited thereto. In some embodiments, the molecular signals generated by the system need not correspond to full nucleotide sequencing reads, but may instead comprise compressed, inferred, or proxy representations of transcriptomic state sufficient for downstream normalization and inference. Such representations may be derived from partial sequencing, targeted panels, hybrid capture, signal-based sensing, or othermolecular measurement techniques capable of reflecting platelet-associated transcriptional patterns. The physical system is tightly integrated with the processors described herein (such as the OncoSage platform). Data generated by the device are transmitted to the platform, where they undergo preprocessing (including Stable Ratio Normalization), vectorization, and classification using trained machine learning models. The processor then outputs a diagnostic assessment. For example, the systems and neural network models described herein outputs a decision support interface that presents a visual graphic representation of a patient’s biomarker panel. In another example, the visual graphic representation of a patient’s biomarker panel is dynamically updated based on time point triggers, such as repeated physical sample collection following medical interventions or medical events. Metadata related to the physical sample batch information (i.e. time stamp, associated medical interventions or medical events) is embedded to the visual graphic representation of the patient’s biomarker panel for interactive presentation based on selected batch information. In another example, the neural network model outputs a decision support interface that automatically updates a list available approved drugs based on the type of cancer diagnosed by the system. The decision support interface stores a master list of approved therapies and drugs in a memory storage device. When the system outputs an positive diagnostic assessed of cancer, metadata representing the associated cartridge which was used in the physical processing of the physical biological sample and / or associated RNA sequence structured data from the reader device or a separate RNA sequencing instrument is obtained to dynamically update the mater list of approved therapies and drugs into one or more cancer-specific list of approved therapies and drugs. For example, targeted therapy approved for lung cancer include but are not limited to adagrasib (Krazati), afatinib dimaleate (Gilotrif), alectinib (Alecensa), amivantamab-vmjw (Rybrevant), atezolizumab (Tecentriq), atezolizumab and hyaluronidase-tqjs (Tecentriq Hybreza), bevacizumab (Avastin), binimetinib (Mektovi), brigatinib (Alunbrig), capmatinib hydrochloride (Tabrecta), cemiplimab-rwlc (Libtayo), ceritinib (Zykadia), crizotinib (Xalkori), dabrafenib mesylate (Tafinlar), dacomitinib (Vizimpro), durvalumab (Imfinzi), encorafenib (Braftovi), ensartinib hydrochloride (Ensacove), entrectinib (Rozlytrek), erlotinib hydrochloride (Tarceva), fam-trastuzumab deruxtecan-nxki (Enhertu), gefitinib (Iressa), ipilimumab (Yervoy), lazertinib mesylate hydrate (Lazcluze), lorlatinib (Lorbrena), necitumumab (Portrazza), nivolumab (Opdivo), osimertinib mesylate (Tagrisso), pembrolizumab (Keytruda), pralsetinib (Gavreto), ramucirumab (Cyramza), repotrectinib (Augtyro), selpercatinib (Retevmo), sotorasib (Lumakras), tarlatamab-dlle (Imdelltra), tepotinib hydrochloride (Tepmetko), trametinib dimethyl sulfoxide (Mekinist), tremelimumab-actl (Imjudo), zenocutuzumab-zbco (Bizengri). When system outputs a positive diagnostic assessed of lung cancer, the master list of approved drugs and therapies are updated with a list of targeted therapy approved for lung cancer. For example, the systems and neural network models described herein outputs structured data or metadata(including for example, biomarker profile, cancer diagnostic assessment, cancer likelihood indicator, and / or approved list of targeted therapy) which embedded or appended into patient health records. In some embodiments, the diagnostic assessment comprises, for example, cancer-positive or cancer-negative classification, cancer-type likelihoods, stage or risk stratification, longitudinal trend analysis (based on multiple samplings overtime).

[0180] The system is configured to support repeated testing over time. Each cartridge run contributes to a longitudinal molecular record for a given individual. These records are integrated into a patient-specific digital representation of platelet transcriptomic state over time. Changes in these representations are used to infer disease progression, treatment response, or recurrence. In certain embodiments, the system is configured to operate under variable sample sizes, batch compositions, and control availability. Where control samples are limited or unevenly distributed, the system may adaptively adjust normalization parameters or inference logic while preserving diagnostic validity. Such robustness enables reliable operation across heterogeneous clinical and operational contexts.

[0181] In some embodiments, the system maintains versioned models, preprocessing parameters, and data lineage records to support auditability, regulatory review, and longitudinal consistency. Outputs generated by the system may be traceable to specific device runs, cartridge identifiers, and computational model versions.

[0182] The disclosed embodiments collectively define an end-to-end diagnostic system in which physical sample collection, platelet enrichment, molecular stabilization, data generation, and computational inference are co-designed and interdependent. The biological, mechanical, chemical, and computational components of the system operate in concert such that removal, substitution, or decoupling of any component materially degrades diagnostic performance. Accordingly, the invention is not limited to any single step or module in isolation but rather encompasses the integrated operation of the system as a whole.

[0183] The system described herein may be deployed in a variety of operational configurations, including centralized laboratory settings, decentralized clinical environments, pharmacy-based testing locations, and home-based testing workflows coupled to remote computational analysis. The modular design of the physical devices, consumable cartridges, and computational platform enables flexible deployment without alteration of the underlying diagnostic principles, thereby supporting diverse clinical, regulatory, and commercial use cases.

[0184] Stable Ratio Normalization (StaR-norm) for cancer detection from transcriptomic read data - Transcriptome-based cancer detection methods typically rely on global normalization strategies such as counts per million (CPM), transcripts per million (TPM), trimmed mean of M-values (TMM), quantile normalization, or batch harmonization techniques such as ComBat. These approaches were originally designed for differential expression analysis in relatively homogeneous experimental settings, where most genes are assumed to be similarlyexpressed across samples and only a small subset is differentially regulated. In the context of cancer detection from blood-derived or platelet-derived RNA, these assumptions are systematically violated. Cancer induces systemic, heterogeneous, and non-uniform transcriptional reprogramming, often affecting a large fraction of the measured transcriptome. As a result, A substantial proportion of genes may shift in expression due to tumor-host interactions, inflammation, immune signaling, or platelet education, invalidating the assumption that most genes are stable across samples. Methods such as ComBat attempt to remove batch effects by shrinking expression distributions across batches. In cancer detection tasks, this can inadvertently remove true biological differences when cancer prevalence, subtype composition, or disease severity differs across batches, resulting in distorted biological signal. Normalization parameters learned on one cohort often fail when applied to new cohorts with different sequencing depth, control composition, or biological background, leading to unstable inference. Hence frozen scalers and global transforms do not generalize. Many normalization methods also reduce variance indiscriminately, suppressing weak but informative cancer-associated signals that may only be present in subsets of patients, and therefore cancer heterogeneity is masked rather than revealed. Consequently, models trained using conventional normalization pipelines often exhibit reduced robustness, poor external generalization, and inconsistent performance across cohorts.

[0185] The present inventor have specifically developed a Stable Ratio Normalization (StaR-norm) approach to address the challenges of cancer detection from transcriptom ic read data by reframing normalization as a within-batch, control-anchored ratio problem, rather than a global distribution-matching problem. The core insight underlying StaR-norm approach is that relative expression differences between cancer samples and contemporaneous controls within the same batch are more stable, biologically meaningful, and generalizable than absolute expression levels normalized across unrelated samples or cohorts.

[0186] In certain embodiments, StaR-norm inference is applied to new samples in batches that include both control and cancer samples. In some embodiments, applying the normalization inference comprises the steps of 1) ingesting raw gene-level read counts are for a batch of samples, 2) applying a predefined gene filter (learned during training), 3) adjusting read counts to avoid zero inflation, 4) performing library size normalization using CPM, 5) for each batch, computing the mean expression of control samples per gene, 6) normalizing each sample’s expression as a log-ratio relative to the batch-specific control mean, and 7) aligning the resulting normalized matrix to the training gene space. Notably, this inference process does not use batch harmonization (e.g., ComBat) and does not apply frozen scalers, ensuring that biological signal is preserved.

[0187] Definitions.• Xg sdenote the raw read count for gene g in sample s,• G denote the set of genes retained after filtering,• B denote a batch,• C(B) c B denote the set of control samples within batch B.

[0188] Zero-offset adjustment. To avoid undefined ratios and logarithms, a constant offset is added:Xq s= Xq s+ 1

[0189] Library size normalization (CPM). or each sample s, the library size is computed as:Counts per million are then calculated as:Xq sCPMq s= x 106'SThis step removes variability due to sequencing depth while preserving relative expression structure.

[0190] Batch-specific control reference construction. For each batch B and gene g, the control reference expression is computed as the mean CPM across control samples:This reference represents the baseline transcriptional state for gene g in batch B.

[0191] Stable ratio transformation. Each sample is normalized by computing the log-ratio of its CPM value relative to the batch-specific control mean:>This transformation yields features that represent relative deviations from normal expression, rather than absolute abundance. StaR-norm reduces unwanted variance through several mechanisms. Any batch-level multiplicative biases affecting both cancer and control samples (e.g., sequencing chemistry, RNA quality) are canceled out in the ratio. System-wide expression changes that affect all samples in a batch are normalized away, preventing spurious signal inflation. By expressing values as log-ratios relative to controls, gene-specific dispersion is reduced while preserving directional changes. Unlike variance-shrinking methods, StaR-norm does not force expression distributions to match across batches, avoiding over-correction.

[0192] A key advantage of StaR-norm is its ability to preserve and amplify heterogeneous cancer signals, which are often diluted by global normalization approaches. Because each cancer sample is compared directly to the contemporaneous control baseline: genes dysregulated only in subsets of cancer patients remain detectable, weak but consistent deviations across small patient groups are preserved, and subtype-specific and stage-specificsignals are retained. This property is particularly important for early-stage cancer detection, where transcriptomic changes may be subtle, sparse, and heterogeneous. StaR-norm differs fundamentally from existing methods in that it: 1) uses explicit control anchoring, rather than implicit global assumptions. 2) operates within batches, rather than across unrelated cohorts, 3) avoids batch harmonization techniques that can remove biological signal, 4) produces interpretable log-ratio features aligned with biological deviation, and 5) generalizes robustly to new datasets that include internal controls. These properties result in improved stability, generalization, and diagnostic performance in machine-learning models trained for cancer detection. While described here in the context of cancer detection from platelet or blood-derived RNA, StaR-norm is applicable to any transcriptomic setting in which systemic biological perturbations affect large fractions of the transcriptome, and contemporaneous control samples are available within batches. The method is cancer-agnostic and extensible to multi-cancer detection, longitudinal monitoring, and minimal residual disease (MRD) applications.

Claims

We Claim / 1 claim:

1. A computer system for machine learning based diagnostics, the computer system comprising:a computer processor operating in conjunction with computer memory and non-transitory computer readable data storage,the processor configured to:instantiate a machine learning meta-model including a stacked ensemble of a plurality of base models;train the machine learning meta-model using a population dataset to optimize a loss function adapted to classify a disease state;extract, from a physical platelet sample of tumor educated platelets, RNA associated with a new sample under test;extract, from the RNA associated with the new sample under test, a sequencing dataset associated with the new sample under test, the sequencing dataset extracted in the form of vectorized data inputs;operate, the trained machine learning meta-model in an inference mode to process the vectorized data inputs using the trained machine learning meta-model to generate an output dataset including one or more output logits representative of a classified disease state corresponding to the new sample under test, andoutputs a data object representing cancer assessment identifier for automatic embedding into patient health record.

2. The computer system of claim 1, wherein the stacked ensemble includes at least a treebased learner machine learning model data architecture and a neural network model data architecture configured for operation as a meta-learning pipeline, the tree-based learner machine learning model data architecture having a higher level of dimensionality than the neural network model data architecture.

3. The computer system of claim 1 , wherein the one or more output logits representative of the classified disease state corresponding to the new sample under test are utilized by a treatment workflow engine to automatically generate data objects representing requisitions for additional samples for analysis.

4. The computer system of claim 3, wherein the plurality of base models of the stacked ensemble includes at least one model that is trained based on disease trajectory, the at least one model trained based on disease trajectory being provided historical inputs associated with an underlying individual associated with the new sample under test and the requisitions for additional samples for analysis.

5. The computer system of claim 4, wherein the at least one model trained based on disease trajectory is configured to maintain a digital twin representation of the underlying individual, and the stacked ensemble can be operated in inference to extrapolate a future state classified disease state at a designated future timepoint.

6. The computer system of claim 5, wherein the digital twin representation includes one or more data fields corresponding to changed gene and molecular pathways that have changed between different samples under test.

7. The computer system of claim 6, wherein the changed gene and molecular pathways that have changed between different samples under test are used to identify potential cross-talk linkages between the changed gene and molecular pathways.

8. The computer system of claim 2, wherein the level of dimensionality for each of the base models of the stacked ensemble is controlled by one or more controllable parameters.

9. The computer system of claim 8, wherein the one or more controllable parameters are controlled based on data inputs indicative of available computing resources of the computer system.

10. The computer system of claim 1, wherein the population dataset comprises a first subpopulation of healthy patients, a second subpopulation of target cancer patients, a third subpopulation of non-target cancer patients.

11. The computer system of claim 1 , comprising a physical input receptacle for receiving the physical platelet sample of tumor educated platelets, a cartridge for outputting a platelet- enriched fraction derived from the physical platelet sample, the cartridge comprising one or more of:i. a microfluidic channel for purification of tumor educated platelets from the physical platelet sample,ii. gradient chamber for enrichment of tumor educated platelets,iii. one or more buffers for platelet preservation,iv. one or more buffers for stabilizing RNA,v. a surface coated with antibodies to a target biomarker, andvi. a surface for capturing target RNA molecules; andand a reader device for coupling with the cartridge, the reader device being in electronic communication with the processor.

12. The computer system of claim 11 , wherein the physical platelet sample is whole blood.

13. The computer system of claim 11 , wherein the reader device outputs a preprocessed structured dataset representing the new sample under test.

14. The computer system of claim 13, where in the preprocessed structured dataset comprise:quantitative representation of RNA abundance,• gene- or pathway-level signatures,• quality metrics, and / or• batch identifiers.

15. The computer system of claim 13, wherein the processor is configured to correlate the preprocessed structured dataset with the sequencing dataset associated with the new sample under test.

16. A computer method for machine learning based diagnostics, the computer method comprising:instantiating a machine learning meta-model including a stacked ensemble of a plurality of base models;training the machine learning meta-model using a population dataset to optimize a loss function adapted to classify a disease state;extracting, from a physical platelet sample of tumor educated platelets, RNA associated with a new sample under test;extracting, from the RNA associated with the new sample under test, a sequencing dataset associated with the new sample under test, the sequencing dataset extracted in the form of vectorized data inputs;operating, the trained machine learning meta-model in an inference mode to process the vectorized data inputs using the trained machine learning meta-model to generate an output dataset including one or more output logits representative of a classified disease state corresponding to the new sample under test; andoutputting a data object representing cancer assessment identifier and automatically embedding said data object into patient health record.

17. The computer method of claim 16, wherein the stacked ensemble includes at least a tree-based learner machine learning model data architecture and a neural network model data architecture configured for operation as a meta-learning pipeline, the tree-based learner machine learning model data architecture having a higher level of dimensionality than the neural network model data architecture.

18. The computer method of claim 16, wherein the one or more output logits representative of the classified disease state corresponding to the new sample under test are utilized by a treatment workflow engine to automatically generate data objects representing requisitions for additional samples for analysis.

19. The computer method of claim 18, wherein the plurality of base models of the stacked ensemble includes at least one model that is trained based on disease trajectory, the at least one model trained based on disease trajectory being provided historical inputs associated with an underlying individual associated with the new sample under test and the requisitions for additional samples for analysis.

20. The computer method of claim 19, wherein the at least one model trained based on disease trajectory is configured to maintain a digital twin representation of the underlying individual, and the stacked ensemble can be operated in inference to extrapolate a future state classified disease state at a designated future timepoint.

21. The computer method of claim 20, wherein the digital twin representation includes one or more data fields corresponding to changed gene and molecular pathways that have changed between different samples under test.

22. The computer method of claim 21, wherein the changed gene and molecular pathways that have changed between different samples under test are used to identify potential crosstalk linkages between the changed gene and molecular pathways.

23. The computer method of claim 17, wherein the level of dimensionality for each of the base models of the stacked ensemble is controlled by one or more controllable parameters.

24. The computer method of claim 23, wherein the one or more controllable parameters are controlled based on data inputs indicative of available computing resources.

25. The computer method of claim 16, wherein the population dataset comprises a first subpopulation of healthy patients, a second subpopulation of target cancer patients, a third subpopulation of non-target cancer patients.

26. The computer method of claim 16, comprising preprocessing parameters of the physical platelet sample of tumor educated platelets into a preprocessed structured dataset representing the new sample under test.

27. The computer method of claim 26, where in the structured data comprise:• quantitative representation of RNA abundance,• gene- or pathway-level signatures,• quality metrics, and / or• batch identifiers.

28. The computer method of claim 26, comprising correlating the preprocessed structured dataset with the sequencing dataset associated with the new sample under test.

29. The computer method of claim 16, comprising normalizing the sequencing dataset by normalizing each sample expression as a log-ratio relative to a batch-specific control mean value.

30. The computer method of claim 16, comprising generating a cancer diagnostic assessment based on the output dataset, the cancer diagnostic assessment comprising cancer classification, risk stratification, trend analysis.

31. The computer method of claim 30, further comprising identifying a drug treatment regime based on the cancer diagnostic assessment.

32. A non-transitory computer readable medium storing machine interpretable instructions, which when executed by a processor, cause the processor to execute a method for machine learning based diagnostics according to any one of claims 16-31.