Method and system for determining her2 status using molecular data
A machine learning-based method processes molecular data to accurately predict HER2 status, addressing the inefficiencies and inconsistencies of existing methods and improving treatment decisions for HER2-low tumors.
Patent Information
- Application Number
- JP2024199275
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-15
- Filing Date
- 2024-11-14
- Publication Date
- 2025-05-27
AI Technical Summary
Current methods for determining HER2 status, such as IHC and FISH, are inefficient and inconsistent, leading to inaccurate classification of HER2-low tumors and inadequate treatment decisions.
A computer-implemented method using a multi-stage machine learning architecture to process molecular data and predict HER2 status, including training models to distinguish between HER2-positive, HER2-low, and HER2-negative statuses.
This approach enhances the accuracy and reproducibility of HER2 status testing, improves the identification of HER2-low patients who may benefit from targeted therapies, and reduces the need for invasive and costly conventional tests.
Smart Images

Figure 2025081279000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims priority to U.S. Patent Application No. 63 / 599,508, filed on November 15, 2023, entitled "METHODS AND SYSTEMS FOR DETERMINING HER2 STATUS USING MOLECULAR DATA", which is hereby incorporated by reference in its entirety.
[0002] The present disclosure is directed to methods and systems for determining HER2 status using molecular data, and more specifically, to techniques for training and operating one or more machine - learning models for processing a patient's molecular data to predict HER2 status, including HER2 low status.
Background Art
[0003] The description of the background art provided herein is for the purpose of generally presenting the context of the present disclosure. The research of the inventors named as such at present, to the extent described in this background art section, and aspects of the description that may not be considered prior art at the time of filing in another way, are not admitted as prior art to the present disclosure, either expressly or implicitly.
[0004] Human epidermal growth factor receptor 2 (HER2) is a protein that can play a significant role in certain types of cancer (e.g., breast cancer). HER2 is a receptor protein found on the surface of some normal cells and can be overexpressed in cancer cells. When HER2 is overexpressed, it can lead to uncontrolled cell growth and the development of invasive tumors. A patient's HER2 status is an important factor in cancer diagnosis and treatment. To determine the HER2 status, tumors have conventionally been classified into one of the following groups. 1. HER2 negative (HER2-): This group includes cancer tumors in which the HER2 protein is not overexpressed or the HER2 expression level is very low. HER2-negative cancers are not usually treated with drugs that specifically target HER2. 2. HER2 positive (HER2+): These tumors have overexpression of the HER2 protein, and this overexpression contributes to the rapid growth and invasiveness of the cancer. HER2-positive cancers are likely to respond to targeted therapies that block the HER2 receptor, such as trastuzumab (Herceptin) and other HER2-targeted drugs. These targeted therapies can be used in addition to standard treatments such as chemotherapy, surgery, and radiation therapy. 3. HER2 equivocal: In some cases, the HER2 expression level in a tumor can be borderline or ambiguous, meaning it is not clearly positive or negative. In such cases, additional tests such as fluorescence in situ hybridization (FISH), chromogenic in situ hybridization (CISH), and / or immunohistochemistry (IHC) may be performed to more definitively determine the HER2 status. Generally, FISH and IHC are the standard methods conventionally used to determine the HER2 status.
[0005] More recently, a new classification called HER2-low (sometimes referred to as HER2-low or HER2-2+) has been developed to characterize the HER2 status of patients. This category is used to describe tumors with lower HER2 expression levels, unlike HER2-negative and HER2-positive cancers. HER2-low cancers are characterized by having HER2 expression that falls within an intermediate range, meaning that the HER2 expression is higher than that of HER2-negative (HER2-) but not high enough to be classified as HER2-positive (HER2+). HER2-low tumors typically have low to moderate HER2 protein expression and are generally not amplified. The classification of HER2-low breast cancer has gained importance because it has been identified as a distinct subgroup with its own characteristics and potential therapeutic implications. Since the importance of the HER2-low category has not been widely recognized until now, conventional HER2 tests have not been designed to accurately classify patients into the HER2-low category. HER2-low tumors do not respond very well to conventional HER2-targeted therapies such as trastuzumab (Herceptin), which is effective in HER2-positive breast cancer, but recent studies have shown that some experimental targeted therapies may be beneficial for HER2-low breast cancer.
[0006] One such experimental therapy is trastuzumab deruxtecan (T-DXd), which has been shown to be promising in clinical trials for treating HER2-low breast cancer. This drug is an antibody-drug conjugate that specifically targets HER2 and delivers a chemotherapy drug to cancer cells, making it a potential treatment option for HER2-low breast cancer.
[0007] However, there is currently no effective means to identify patients who will benefit from the drug and have HER2 protein expression. At first glance, IHC and FISH seem like candidates for detecting the HER2-low state, but IHC and FISH have several technical drawbacks. First, IHC and FISH require multiple tissue slides (one and three or more, respectively). Often, the slide material (e.g., tumor blocks) available for performing IHC and / or FISH in addition to (or instead of) other uses of the tissue is not sufficient. Furthermore, the assays using IHC and FISH are protein-based and generally, a single slide must be used for each protein. The input to IHC and FISH is generally formalin-fixed paraffin-embedded (FFPE) stained tumor block slides. Thus, usually, the number of microscope slides determines the number of tests that can be performed. For example, one slide is required to perform IHC, and more slides are required for FISH.
[0008] Furthermore, IHC requires expensive antibodies. Staining takes time, the pathological review is usually performed manually, and the assessment takes a significant amount of time.
[0009] Moreover, the accuracy and consistency of IHC / FISH interpretation are lacking. For example, one pathologist may annotate a slide as HER2-2+, while another pathologist may annotate the slide as HER2-1+. Here, generally, IHC0 is HER2 negative, IHC1+ is HER2 low, IHC2+ and FISH negative is HER2 low, IHC2+ and FISH positive is HER2 positive, and IHC3+ is HER2 positive. In such cases, the results are not useful for downstream processes. Overall, given the high level of disagreement among pathologists regarding the HER2 IHC status, particularly regarding the determination of the HER2-low state (and especially the distinction between the HER2-low state and the HER2-negative state respectively), it has been reported that there is a high variability in the determination of the HER2-low state via IHC.
[0010] Inconsistent annotations can lead to discrepancies in interpretation and diagnosis, potentially affecting disease diagnosis and classification, and thus having an adverse impact on clinical practice in the fields of pathology, oncology, and disease diagnosis. For example, in cancer diagnosis, inconsistent annotations of tumor markers can lead to misclassification and result in incorrect treatment decisions. Inconsistent annotations can also lead to under-treatment or over-treatment, potentially affecting patient outcomes and the quality of care. Inconsistent annotations in research can compromise the reproducibility of scientific findings and prevent the replication or validation of research results. Inconsistent annotations can have an adverse impact in the areas of data quality, biomarker discovery, quality control, and resource efficiency.
[0011] Inconsistent annotations can also be caused by simple confusion in numerical scoring in different ways of labeling the HER2 status. In the past, pathologists were trained to consider IHC scores of 0 and 1 as equivalent to HER2 negative. The number 2 indicated HER2 equivocal, and if subsequent FISH testing was positive, these results were labeled HER2 positive, and if subsequent FISH testing was negative, these results were labeled HER2 negative. The number 3 was classified as HER2 positive. In view of the fourth category (HER2 low), 0 is interpreted as HER2 negative, 1 is interpreted as HER2 low, and 2 and 3 are still classified in the same way (IHC2+, FISH negative is HER2 low, IHC2+, FISH positive is HER2 positive). Having two systems in which numerical scores can refer to different meanings has caused confusion among the pathologist community.
[0012] Furthermore, since IHC and FISH tests are not routinely used to assess the HER2 status of all cancers, these values are not always available for analysis. For example, in some gastrointestinal cancers, IHC and FISH are not usually part of the clinical workflow. Therefore, the data may not be available for modeling. Metastatic breast cancer spreading to the gastrointestinal tract may not be tested for HER2 because clinicians are unaware that this metastatic cancer is breast cancer. Many patients may not have IHC results even if they have RNA / DNA samples. Furthermore, FISH is much more expensive than IHC because more slides are required to perform FISH.
[0013] More accurately determining which patients are likely to respond to HER2-targeted therapy is important as it helps guide treatment decisions for cancer patients. For example, HER2-targeted therapy has been shown to be highly effective in HER2-positive breast cancer as determined by IHC / FISH, improving both survival rates and treatment outcomes. In contrast, HER2-negative breast cancer patients are typically treated with other therapies and HER2-targeted drugs are not used as they are unlikely to be beneficial. However, recent studies have shown that HER2-low patients, previously identified as a subset of the HER2-negative population, may also benefit from HER2-targeted therapy, indicating the need for a more accurate method to distinguish HER2-low patients from those without HER2 protein expression.
[0014] Therefore, there is an opportunity to improve platforms and technologies for determining HER2 status using molecular data by enhancing the reproducibility, accuracy, and sensitivity of HER2 status testing. SUMMARY OF THE INVENTION
[0015] In one aspect, a computer-implemented method for determining a patient's HER2-low status using the patient's molecular data includes: (a) receiving digital biological data via one or more processors; (b) processing the digital biological data corresponding to the patient using a trained multi-stage machine learning architecture via one or more processors, the processing including: (i) processing the digital biological data using a trained HER2-positive model to determine whether the digital biological data indicates that the patient's HER2 status is HER2-positive; (ii) processing the digital biological data using a trained HER2-low model to identify whether the patient's HER2 status is HER2-low when the patient's HER2 status is not HER2-positive; (iii) designating the patient's HER2 status as HER2-negative when the patient's HER2 status is neither HER2-positive nor HER2-low; (c) generating a digital HER2-low status report corresponding to the patient via one or more processors; and (d) displaying the digital HER2-low status report via a display device.
[0016] In another aspect, a computer-implemented method for training a model architecture for determining a patient's HER2-low status using the patient's molecular data includes: (a) receiving training digital biological data including a plurality of molecular signatures each having a respective label via one or more processors; (b) initializing a machine learning model having a plurality of hyperparameters in a computer's memory via one or more processors; (c) processing, via one or more processors, the plurality of molecular signatures and the respective labels of the plurality of molecular signatures in the training digital biological data using the machine learning model to generate a trained machine learning model; and (d) storing the trained machine learning model in the computer's memory via one or more processors, the storing including generating a serialized copy of the machine learning model and writing the serialized copy of the machine learning model to the computer's memory.
[0017] In yet another aspect, a computing system includes one or more processors and one or more memories storing computer-readable instructions that, when executed, cause the computing system to: (a) receive digital biometric data; (b) process the digital biometric data corresponding to a patient using a trained multi-stage machine learning architecture, the processing including: (i) processing the digital data using a trained HER2 positive model to determine whether the digital biometric data indicates that the patient's HER2 status is HER2 positive; (ii) processing the digital biometric data using a trained HER2 low model to identify whether the patient's HER2 status is HER2 low when the patient is not HER2 positive; and (iii) designating the patient's HER2 status as HER2 negative when the patient is neither HER2 positive nor HER2 low; and (c) generate a digital HER2 low status report corresponding to the patient via one or more processors; and (d) cause the digital HER2 low status report to be displayed via a display device.
[0018] In yet another aspect, a computer-readable medium includes computer-executable instructions that, when executed, cause a computer to: (a) receive digital biometric data; (b) process the digital biometric data corresponding to a patient using a trained multi-stage machine learning architecture, the processing including: (i) processing the digital data using a trained HER2 positive model to determine whether the digital biometric data indicates that the patient's HER2 status is HER2 positive; (ii) processing the digital biometric data using a trained HER2 low model to identify whether the patient's HER2 status is HER2 low when the patient is not HER2 positive; and (iii) designating the patient's HER2 status as HER2 negative when the patient is neither HER2 positive nor HER2 low; and (c) generate a digital HER2 low status report corresponding to the patient via one or more processors; and (d) cause the digital HER2 low status report to be displayed via a display device.
[0019] In a further aspect, a computing system includes one or more processors and one or more memories storing computer-readable instructions that, when executed, cause the computing system to: (a) receive training digital biometric data including a plurality of molecular signatures each having a respective label; (b) initialize a machine learning model having a plurality of hyperparameters in the memory of the computer; (c) process the plurality of molecular signatures and the respective label of each of the plurality of molecular signatures in the training digital biometric data using the machine learning model to generate a trained machine learning model; and (d) store the trained machine learning model in the memory of the computer, the storing including generating a serialized copy of the machine learning model and writing the serialized copy of the machine learning model to the memory of the computer.
[0020] In yet another aspect, a computer-readable medium includes computer-executable instructions that, when executed, cause a computer to: (a) receive training digital biological data including a plurality of molecular signatures each having a respective label; (b) initialize a machine learning model having a plurality of hyperparameters in a memory of the computer; (c) process, using the machine learning model, the plurality of molecular signatures and the label of each of the plurality of molecular signatures in the training digital biological data to generate a trained machine learning model; and (d) store the trained machine learning model in a memory of the computer, the storing including generating a serialized copy of the machine learning model and writing the serialized copy of the machine learning model to the memory of the computer. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The drawings described below depict various aspects of the systems and methods disclosed herein. It should be understood that each drawing depicts an example of an aspect of the present system and method.
[0022]
Figure 1
Figure 2
Figure 3A
Figure 3B
Figure 3C
Figure 4
Figure 5
[0023] Overview The present technique is directed to methods and systems for determining HER2 status using a patient's molecular data, and more specifically, to methods and systems for training and operating one or more models for processing a patient's molecular data to predict the HER2 status (as defined by IHC / FISH testing). The present technique enables a clinician to receive a prediction score that a given patient will have a HER2 low status when evaluated using IHC and FISH. This advantageously enables the clinician to obtain results equivalent to IHC and FISH testing without actually performing those tests, saving on material costs, avoiding potential inaccuracies of those tests, or warning the clinician that patients who would not normally undergo a HER2 test may potentially benefit from a HER2 test, as described above.
[0024] In some embodiments, the present technique may generate a molecular signature using one or more portions of a patient's molecular data. For example, copy number variant (CNV) data and / or RNA data may be joined via concatenation or otherwise to form a molecular signature. The present technique may, in some embodiments, generate a molecular signature using additional / different data. For example, the features included in the molecular signature may include DNA data in some embodiments.
[0025] As described, IHC and FISH can lead to inconsistently annotated slides and may not be usable when tumor blocks are insufficient. Thus, the advantage of using molecular assays such as the present technique is that, when RNASeq has already been performed, predictive factors can be run on RNASeq data and / or DNASeq data to predict the HER2 status regardless of whether additional tissue is available. As noted above, not only may there be insufficient tissue, but these values are not always available for analysis because IHC and FISH assays are not routinely used to assess the HER2 status of all cancers. Many patients may have RNA / DNA samples but not IHC results, and the present technique advantageously expands the population of patients who may be eligible for HER2-based targeted therapies. Thus, the present technique has an enhanced ability compared to conventional techniques. Another advantage of this molecular assay is that the problem of inconsistent annotation is solved. Further, the technique can process all 20,000 genes at once using RNASeq, which avoids the bottlenecks of IHC and FISH, where each run of the assay is generally limited to fewer than 50 genes.
[0026] As considered, the FISH assay is much more expensive than the IHC assay (more cost, more slides are required to perform FISH). Thus, the present technique is particularly compelling for IHC2+ cases. In future aspects, when modeling more proteins beyond HER2, the present technique does not require additional tumor slides to re-run the IHC assay for different proteins (e.g., TROP2).
[0027] This technique may include a feedback mechanism, whereby drug response data of one or more patients whose molecular data is examined to determine HER2 status prediction is fed back into the model training pipeline, enabling the model to learn better criteria for predicting the HER2 status of patients. By having molecular modeling with feedback, it can help in fine-tuning the threshold by reconciling the ground truth from future results with past predictions. Specifically, RNA-seq training data can be augmented with data indicating whether a patient responded to HER2 targeted therapy (whether the patient was a responder or not), which can be used as a training target instead of the HER2 status determined by IHC / FISH.
[0028] In some embodiments, the technique can obtain sequencing data (e.g., by scraping) using slides in the same manner as performing IHC and FISH. In some embodiments, the use of slides can be completely skipped or omitted. For example, biological patient data can be obtained directly from a surgical intake or biopsy, and nucleic acids can be isolated directly from the material. The technique does not require staining. Tissues are broken down by a liquid buffer and enzymes, enabling RNA and / or DNA to be isolated and prepared for a sequencing machine that generates RNASeq and / or DNAseq data. CNVs can be determined from DNA data and / or inferred from RNA data (e.g., amplified / increased CNVs can be associated with high RNA expression levels, and loss CNVs can be associated with low RNA expression levels. Generally, CNV data represents deletions or extra copies of chromosomes within a cell, as seen in cancer cells. Since HER2 positivity may be associated with an increased copy number of the HER2 gene, RNASeq and CNV data can be used to predict the HER2 status. Thus, since slides are not consumed or used as in the case of conventional IHC or FISH tests, the technique is less destructive than conventional techniques. The technique is also advantageous because it does not require the cost of antibodies, staining, test slides, and other materials. Furthermore, the technique does not require a human to be "in the loop" of a predictive algorithm as IHC and FISH often do. Moreover, as discussed above, the same RNAseq and DNAseq data (or derivatives such as CNV data) processed by this modeling will give the same answer (or a modifiable different answer) when run sequentially, while the analysis of a human pathologist may not, making it more deterministic to the extent possible. Tumor blocks do not necessarily give exactly the same RNAseq results (due to machine variation). Usually, these variations do not affect the modeling output unless the sample is close to the threshold. However, a particular set of RNAseq / DNAseq data will always yield the same output from current modeling techniques.
[0029] In particular, since both the 0 and 1 HER2 states were previously classified as HER2 negative, pathologists may not have paid much attention to the distinction between the IHC scores of 0 and 1. Another advantage of this technique is that RNASeq analysis is not affected by human subjectivity, and even if patients in the training data are mis-scored as 0 when they should actually be 1+ (or vice versa), this machine learning-based technique can accurately classify the patients.
[0030] In some examples, CNV data alone can distinguish between HER2 positive and non-HER2 positive well, but has a significant dynamic range. This multimodal technique that uses both CNV and RNASeq data to predict HER2 status enables more accurate prediction of HER2 positive and HER2 low subtypes.
[0031] Furthermore, in some aspects, data of patients previously assigned the HER2 negative state can be processed using the second part (the second stage) of the model architecture to identify patients who are HER2 low. Thus, in some aspects, if another technique is already being used to identify a patient's HER2 positive / HER2 negative state, or if HER2 positive is not of interest for other reasons, only the second part of the architecture can be used. In this example, any HER2 positive included in the data provided to the stage 2 model can be classified as HER2 low. In this case, HER2 low may not be an accurate label, and another label can be used (e.g., responder / non-responder - the sample refers to a patient who is either HER2 positive or HER2 low).
[0032] In some cases, there is no perfect correlation between the IHC / FISH status and the drug response. In some embodiments, instead of using IHC / FISH for training, drug response data can be used for training. After the drug response data becomes available, considering the discrepancies between the IHC / FISH labels and the molecular subtypes and response data can provide more clues. Questions that can be answered include: 1) whether there are IHC / FISH subtypes that better explain or predict the drug response, 2) whether there are molecular subtypes that better explain or predict the drug response, and / or 3) whether there is a cascade of IHC / FISH labels and molecular subtypes that better explain or predict the drug response.
[0033] Exemplary computer-implemented machine learning model Figure 2 depicts an exemplary multi-stage machine learning model architecture 200 according to some embodiments. In some embodiments, the model architecture 200 (e.g., a prediction / diagnosis model having two stages) can be trained in a feature space composed of only RNA features. In some embodiments, the model architecture 200 (e.g., a multi-omics prediction / diagnosis model having two stages 202) can be based on a feature space composed of RNA features and copy number variant features. Each of the two stages 202 can include respective model types (e.g., respective random forest models trained with CNV- and RNA-based features), but the models of each stage 202 can be trained with different labels. Specifically, the first stage model 202a can be trained using labeled data having a binary label of (HER2 positive, non-HER2 positive). The second stage model 202b can be trained using labeled data having a binary label of (HER2 low, non-HER2 low).
[0034] For example, random forest model training can be implemented using libraries such as Scikit - learn, R Random Forest, Java Weka, etc. The RNA features used for training may correspond to one or more of the genes "GRB7", "GSDMB", "MIEN1", "ORMDL3", "PGAP3", "PSMD3", "STARD3", and / or "ERBB2". As contemplated, in some aspects, copy number variant scores or counts can be used as training features. Multiple copy number variant features across multiple genes can be used for training. For example, the copy number variant features used for training may correspond to one or more of the genes "BRD4", "CDK12", "ERBB2", and / or "RARA". In some aspects, more and / or different features can be used for training.
[0035] For RNA features, the technique can generate feature values for training that are the log of transcripts per million of the RNA expression of the gene. The feature values can be continuous non - negative numbers. The corresponding training values can be the HER2 labels discussed above. The training labels can be determined using IHC and / or FISH. Thus, a model trained using this training data (e.g., architecture 200) can learn to predict an output that is consistent with the IHC and FISH techniques for similar samples. The IHC label values can be 0, 1, 2, or 3. The value of the FISH label can be positive or negative. Generally, an IHC label value of 0 corresponds to a HER2 - negative sample. An IHC label value of 1 corresponds to a HER2 - low sample, and an IHC label value of 2 requires the use of FISH to resolve further ambiguity. In this case, if FISH is positive, the sample is HER2 - positive. If FISH is negative, the sample is HER2 - low. If the IHC label value is 3, it indicates that the sample is HER2 - positive. Training the machine - learning model herein can include fitting RNA expression values to IHC / FISH label data using a strategy that includes a random forest algorithm.
[0036] Regarding copy number variation training data, the present technique may use some algorithm to generate a continuous number representing the estimated number of copy number variations of a given sample for each gene. For example, the method described in U.S. Patent No. 11,715,226, entitled "Data based cancer research and treatment systems and methods", filed on October 18, 2019, which is hereby incorporated by reference in its entirety for all purposes, may be used to generate an estimate of the number of copy number variations. Specifically, as described in the '226 patent, FIG. 313 illustrates a CNV pipeline that may be used for this purpose. In some cases, two features: a major copy number feature and a subsequent copy number feature may be used to train the architecture 200. The major copy number feature may estimate the copy number of a chromosome. The subsequent copy number feature may be an estimated value of the confidence in the major copy number. These two features may be used for each of the copy number variation genes used for training.
[0037] During operation, the model architecture 200 (i.e., the stage 1 model 202a and the stage 2 model 202b) may sequentially assign labels to samples in a cascade. In the first stage, the model 202a identifies HER2 positive samples (block 204a), and in the second stage, the model 202b identifies HER2 low samples (block 204b). The architecture 200 may annotate samples that are not annotated as either HER2 positive or HER2 low as HER2 negative (block 204c). In each of these two steps, the model 200 may include one or more random forest models trained on RNA-based features and optionally copy number variant features. The model 202 may be trained with different labels. Specifically, the first stage model 202a may be trained with data having (HER2 positive, non-HER2 positive) labels, while the second stage model 202b may be trained data labeled with (HER2 low, non-HER2 low) labels.
[0038] In some embodiments, architecture 200 can be trained using only primary cancer data or metastatic cancer data. In some embodiments, architecture 200 can be trained to make predictions regarding specific cancer types such as breast cancer. In some embodiments, architecture 200 can be trained for a subset of cancer types. Selection of different training strategies / purposes may require different training sets to be constructed. Generally, in empirical tests, models trained with data from samples of approximately 900 patients have shown a high degree of confidence in their predictions.
[0039] It is contemplated that architecture 200 can be trained using algorithms other than random forest. For example, XGBoost, logistic regression, and support vector machines are alternative algorithms to random forest.
[0040] In some embodiments, the technique can be provided to another party via, for example, platform as a service (PaaS) offerings, software as a service (SaaS), white-labeled services, etc. Such other parties can include pharmaceutical companies, biotech companies, contract research organizations, startups (e.g., personalized medicine companies), healthcare analytics companies, software companies, pharmacogenomics companies, healthcare consultants, non-profit research organizations, and the like.
[0041] In some instances, a patient's HER2 status can be utilized to determine whether the patient meets the inclusion or exclusion criteria of a specific clinical trial. Alternatively, the HER2 status of a large number of patients within a database can be employed to estimate the number of patients who may be suitable candidates for a particular trial. This approach can involve processing molecular data, including RNA and potentially DNA data or copy number variant data, to predict a patient's HER2 status. For example, by employing a trained multi-step machine learning architecture, the process can involve determining whether the molecular data indicates a HER2 positive state. If not HER2 positive, the data can be further analyzed to identify whether the patient's status is HER2 low. If the patient's status is neither HER2 positive nor HER2 low, the status can be designated as HER2 negative. This methodology enables the generation of a digital HER2 low status report for a patient and can be presented for the purpose of providing information for clinical trial eligibility determination. Utilizing molecular data for HER2 status prediction not only enhances the accuracy and reproducibility of HER2 status testing but also expands the potential patient population suitable for HER2-based targeted therapies by identifying patients with HER2 low status who may benefit from such treatments.
[0042] Specifically, as contemplated, IHC / FISH may not be clinically performed because the tissue is insufficient, has a low likelihood of being positive, or is stuck in old clinical practice patterns. The machine learning-based molecular diagnostic test of the present invention enables expression calls for any patient receiving a next-generation sequencing report, even if the patient has not undergone an IHC or FISH test. This assay can also be used in pharmaceutical applications for identifying HER2-low and HER2-positive patients using only next-generation sequencing data. The HER2-low patient population has attracted attention in the development of targeted therapies, especially after the success of the Destiny-Breast04 trial. The need to diagnose the ultra-low HER2 population increases if the results of the Destiny-Breast06 trial are positive, and the molecular signature for diagnosing ultra-low HER2 is still useful for finding the minimum HER2 threshold leading to drug (e.g., trastuzumab deruxtecan (TDxd)) response. Generally, given the high variability in HER2-low prediction via IHC / FISH, molecular assays (including the present disclosure) will replace the IHC / FISH assay for HER2-low diagnosis.
[0043] The architecture 200 can be trained to use a selected feature space (i.e., signature) to predict the HER2 status. Generally, the signature includes RNASeq across one or more genomic regions and CNV genomic data. Optionally, a subset of genes can be analyzed.
[0044] In some embodiments, the machine learning architecture 200 may include multiple stages represented by multiple models 202, each making one or more respective predictions 204. The models may use variable screening, future screening, random forest trees, and / or binary classifiers. For example, in a first stage 202a, a predictor model may identify HER2-positive samples (block 204a). In a second stage 202b, a predictor model may identify HER2-low samples (block 204b). Samples that are not HER2-positive or HER2-low may be classified as HER2-negative (block 204c). The model architecture 200 may start without indicating which features are relevant. The model architecture 200 may be trained by a process of eliminating irrelevant features. Feature screening may be suitable for high-dimensional feature spaces to avoid the curse of dimensionality when performing model or variable selection. Thus, variable screening may be used prior to variable selection in this technique.
[0045] In the case of HER2 screening, the identification of HER2-positive samples in the first stage 202a of the model architecture 200 may be performed with a much higher relative confidence than the identification of HER2-low samples in the second stage 202b. This difference in confidence may be due to the fact that HER2-low is a new classification that does not yet have a rigorous clinical gold standard. Thus, the technique may include ruling out results that are likely to be HER2-positive in the first stage 202a, which is performed with a relatively high confidence (e.g., using a binary classifier trained to identify a given sample as corresponding or not corresponding to HER2-positive tissue). In other words, irrelevant variables may be excluded in stage 202a and the overall analysis may be simplified. After detecting HER2-positive at block 204a, the architecture 200 may proceed to the second stage 202b (e.g., using a binary classifier trained to identify a given sample as corresponding or not corresponding to HER2-low tissue).
[0046] In one embodiment, an advantageous advantage of the cascade model architecture 200 (i.e., multi-stage model) is that when HER2-positive data is removed at stage 202a (block 204a), the biological dataset becomes much smaller and easier to manage. At stage 202b, the question is how many of the remaining samples are HER2-low. Stage 202b may include a second trained machine learning model, which annotates the sample data as corresponding (or not corresponding) to HER2-low negative samples and annotates all others as HER2-negative.
[0047] In one embodiment, an important advantage of the present multi-stage cascade model architecture 200 is that the model at the first stage 202a can be trained to detect HER2-negative samples and the second model at the second stage 202b can be trained to detect HER2-low samples, so there is no need to train the model to identify HER2-negative samples. This represents a significant savings in computational effort / work. This is because any given HER2-negative sample can correspond to any of several different types of cancer, including breast cancer HER2-negative cancer samples, primary HER2-negative cancer samples, metastatic HER2-negative cancer samples, colorectal HER2-negative cancer samples, pancreatic HER2-negative cancer samples, bladder HER2-negative cancer samples, etc. The list of potential HER2-negative cancers is long. Since HER2-positive and HER2-low cancer samples are a much smaller class of possible sample results, most of the complexity of screening for HER2-negative cancer samples (a much larger class of potential classification results) is avoided.
[0048] This technique describes a multi-stage model architecture 200 that is essentially optimized for processing RNA Seq data. Due to its cascade structure, it is envisioned that other less efficient architectures may nevertheless be suitable for solving the HER2 classification problem. For example, modifying the training and / or inference data type may result in a more suitable architecture. For example, in some cases, the second stage 202b may be configured as an organ-agnostic diagnostic machine learning model in which a plurality of sub-models perform classification. For example, the second stage 202b may be composed of a model trained for breast cancer, a model trained for pancreatic cancer, a model trained for bladder cancer, and the like. The amount and type of data may determine such architectural decisions. It will also be understood that increasing the number of prediction models may reduce bias but is likely to require more data (i.e., bias-variance trade-off).
[0049] Exemplary Computing Environment FIG. 1 illustrates an exemplary computing environment 100 for implementing the present technique in accordance with some aspects. The environment 100 may include computing resources for determining a patient's HER2 status and / or for further processing / computation based on such predictions, such as patient identification / notification and report generation.
[0050] The computing environment 100 may include a HER2 status prediction computing device 102, a client computing device 104, an electronic network 106, a sequencer system 108, and an electronic database 110. The HER2 status prediction computing device 102 may include an application programming interface 112 that enables programmatic access to the HER2 status prediction computing device 102. The components of the computing environment 100 may be communicatively connected to each other via the electronic network 106 in some aspects. Here, each will be described in more detail.
[0051] The HER2 status prediction computing device 102 can implement, among other things, the training and operation of a machine learning model for predicting the HER2 status of one or more patients, patient identification, and report generation. In some aspects, the HER2 status prediction computing device 102 can be implemented as one or more computing devices (e.g., one or more servers, one or more laptops, one or more mobile computing devices, one or more tablets, one or more wearable devices, one or more cloud computing virtual instances, etc.). The HER2 status prediction computing device 102 can include one or more processors 120, one or more network interface controllers 122, one or more memories 124, an input device 126, and an output device 128.
[0052] In some aspects, the one or more processors 120 can include one or more central processing units, one or more graphics processing units, one or more field programmable gate arrays, one or more application specific integrated circuits, one or more tensor processing units, one or more digital signal processors, one or more neural processing units, one or more RISC-V processors, one or more coprocessors, one or more special processors / accelerators for artificial intelligence or machine learning specific applications, one or more microcontrollers, etc.
[0053] The HER2 status prediction computing device 102 can include one or more network interface controllers 122, such as an Ethernet network interface controller, a wireless network interface controller, etc. The network interface controller 122 can include advanced functions, such as hardware acceleration, special networking protocols, etc., in some aspects.
[0054] The memory 124 of the HER2 status prediction computing device 102 may include volatile and / or non-volatile memory media. For example, the memory 124 may include one or more random access memories, one or more read-only memories, one or more cache memories, one or more hard disk drives, one or more solid state drives, one or more non-volatile memory express, one or more optical drives, one or more universal serial bus flash drives, one or more external hard drives, one or more network-connected storage devices, one or more cloud storage instances, one or more tape drives, etc.
[0055] The memory 124 may store one or more modules 130, for example, as one or more sets of computer-executable instructions. In some embodiments, the module 130 may include additional storage devices such as one or more operating systems (e.g., Microsoft Windows, GNU / Linux, Mac OSX, etc.). The operating system may be configured to execute the module 130 during the operation of the HER2 status prediction computing device 102. For example, the module 130 may include additional modules and / or services for receiving and processing quantitative data. The module 130 may be implemented using any suitable computer programming language (e.g., Python, JavaScript, C, C++, Rust, C#, Swift, Java, Go, LISP, Ruby, Fortran, etc.). The memory may be non-transitory memory.
[0056] The module 130 may include a machine learning model training module 152, a model operation module 154, a patient identification module 156, and a report generation module 158. In some embodiments, more or fewer modules 130 may be included. The modules 130 may be configured to communicate with each other (e.g., via inter-process communication, via a bus, a message queue, a socket, etc.).
[0057] The machine learning model training module 152 may include a set of computer-executable instructions for training one or more machine learning models based on training data. The machine learning model training module 152 may often take input data in the form of a dataset and use it to train a machine learning model. The machine learning model training module 152 may prepare the input data by performing data cleaning, feature engineering, data splitting (into training and validation sets), and handling missing or outlier values. The machine learning model training module 152 may select a machine learning algorithm or model architecture to use for the task at hand. Specifically, the machine learning model training module 152 may include a set of computer-executable instructions for implementing a machine learning training architecture such as the cascade model architecture 200 of FIG. 2. The machine learning model training module 152 may select a random forest algorithm for one or more stages of the model architecture 200.
[0058] The machine learning model training module 152 may include instructions for performing hyperparameter tuning (e.g., settings or configurations of the model that need to be specified before training rather than learned from the data). The machine learning model training module 152 may use grid search or other techniques to specify hyperparameters. Examples of hyperparameters that may be selected for a random forest algorithm include the number of decision trees in the forest, the maximum depth of each individual decision tree in the forest, the minimum number of samples required to split a node in a tree, the minimum number of samples required to be in a leaf node, the maximum number of features to consider for splitting, and the bootstrap parameter.
[0059] The machine learning model training module 152 may include instructions for training a machine learning model selected on training data. The training process may include optimizing the parameters of the model to make predictions. In the case of the random forest algorithm, a voting mechanism (e.g., majority vote) may be used to determine the final prediction. After training, the machine learning model training module may evaluate the performance of the trained model using validation data. In the case of the random forest model, the validation techniques may include out-of-bag validation and cross-validation.
[0060] The machine learning model training module 152 may include instructions for serializing and deserializing the stored model. Thereby, the machine learning model training module 152 can store the trained model as data, reload the model without retraining the model, and use the trained model for prediction.
[0061] The model operation module 154 may include computer-executable instructions for operating one or more trained machine learning models. For example, the model operation module 154 may, in some aspects, access next-generation sequencing data via the sequencer 108. The model operation module 154 may load one or more models trained by the model training module 152. The model operation module 154 may receive raw next-generation sequencing data, preprocess it, apply one or more trained machine learning models, and generate one or more predictions (e.g., HER2 status, specifically, HER2 low status) corresponding to the raw next-generation sequencing data.
[0062] The model operation module 154 can receive next-generation sequencing data including DNA sequencing data (e.g., whole-genome sequencing, exome sequencing), RNA sequencing data (RNAseq), ChIP sequencing data (ChIP-seq), etc. The model operation module 154 may include instructions for receiving and processing data encoded in a plurality of different formats (such as FASTQ, BAM, VCF, etc.).
[0063] The model operation module 154 can perform quality control, read alignment, variant calling, and data normalization.
[0064] The model operation module 154 can perform feature engineering to convert raw sequencing data into features that can be used by one or more machine learning models trained by the machine learning training module 152.
[0065] The model operation module 154 can process the received sequencing data using one or more trained machine models (such as a random forest model, a deep learning model, a support vector machine, etc.) to preprocess the sequencing data. The model operation module 154 can generate model performance statistics such as accuracy, precision, recall, F1 score, or AUC-ROC for classification techniques.
[0066] The patient identification module 156 may include instructions for identifying one or more patients based on the output of the machine learning operation module 154. Specifically, a patient may choose to receive an identification notification based on a specific HER2 status, and in particular, a determination of the HER2-low status. As discussed above, the HER2-low status has been shown to be associated with targeted therapy. However, identifying patients has conventionally been very difficult. By performing an analysis of a patient's specific tumor profile, the present technique can accurately identify patients who may benefit from such targeted therapy without the bias inherent in conventional techniques. This enables precision medicine to be provided to patients in a way that improves the field of machine learning-based cancer therapy. Further, the patient identification module 156 may identify patients who may benefit from a study or clinical trial. Thus, in some aspects, the patient identification module 156 may include patient notification instructions configured to automatically notify patients of their eligibility or matching for a clinical trial or targeted therapy based on the output of the machine learning analysis.
[0067] As used herein, targeted therapies matched by machine learning analysis may include systemic therapy, external beam radiation therapy, surgical therapy, observation therapy, HER2-targeted ADC (antibody-drug conjugate) therapy, and the like.
[0068] In some embodiments, the report generation module 158 can generate a report that includes a prediction regarding the confidence level of the classification of patient data (e.g., the HER2 classification result of the models disclosed herein). For example, these reports can take the form of text documents, digital presentations, word processing documents, and the like. For example, the report generation module 158 can generate one or more reports that include the patient's HER2 score. The report can include a stratified likelihood that the patient has a low HER2 status, as would be measured by IHC and / or FISH. The report generation module 158 can also generate statistics and / or predictions of the progesterone receptor positive (PR+) status and estrogen receptor (ER) based on the output generated by one or more trained models. These results enable clear and practical predictions, particularly by providing clinicians with improvements by providing HER2, PR, and ER biomarkers in a single prediction system. This improves systems that do not include each biomarker, helps clinicians make better treatment decisions, and helps automated systems provide better treatment matching.
[0069] The report generation module 158 can include computer-executable instructions for generating machine-readable results. For example, the client computing device 104 can be accessed by a user to view the results generated by the prediction computing device 102. For example, the user can access a mobile device, laptop device, thin client, etc., embodied as the client computing device 104 to view simulation and confidence scoring results regarding samples whose values have been processed by the prediction computing device 102, and / or reports. Information from the prediction computing device 102 can be transmitted via the network 106 (e.g., for display via the viewer application 180).
[0070] The electronic network 106 can communicatively couple the elements of the environment 100. The network 106 can include a public network such as the Internet, a private network such as a research institution's or enterprise's private network, and / or any combination thereof. The network 106 can include a local area network (LAN), a wide area network (WAN), a cellular network, a satellite network, and / or other network infrastructure, whether wireless or wired.
[0071] In some embodiments, the network 106 can be communicatively coupled to and / or be part of a cloud-based platform (e.g., a cloud computing infrastructure). The network 106 can utilize communication protocols including packet-based and / or datagram-based protocols such as Internet Protocol, Transmission Control Protocol, User Datagram Protocol, and / or other types of protocols. The network 106 can include one or more devices that facilitate network communication and / or form the hardware infrastructure for the network, such as one or more switches, one or more routers, one or more gateways, one or more access points (such as wireless access points), one or more firewalls, one or more base stations, one or more repeaters, one or more backbone devices, etc.
[0072] The sequencer system 108 can include next-generation sequencers, such as RNA sequencing whole exome capture transcriptome assays and DNA sequencing assays, which can be whole genome sequencing or a targeted panel (e.g., a targeted oncology panel using a hybrid capture library preparation).
[0073] The electronic database 110 may include one or more suitable electronic databases for storing and retrieving data, such as relational databases (e.g., MySQL database, Oracle database, Microsoft SQL Server database, PostgreSQL database, etc.). The electronic database 110 may be a NoSQL database such as a key-value store, a graph database, a document store, etc. The electronic database 110 may be an object-oriented database, a hierarchical database, a spatial database, a time-series database, an in-memory database, etc. In some embodiments, some or all of the electronic database 110 may be distributed.
[0074] During operation, one or more sequencer executions may be performed using the sequencer 108 by the company operating / controlling the environment 100 or by another interested party. The results of the sequencer may be received as sequencer data by the HER2 status prediction computing device 102. The HER2 status prediction computing device 102 may preprocess the sequencer data and optionally store some or all of it in the electronic database 110 of FIG. 1. The HER2 status prediction computing device 102 may load one or more trained models, provide the sequencer data as input to the one or more trained models, and receive a prediction regarding the HER2 status from the one or more trained models. In some embodiments, the HER2 status prediction computing device 102 may receive DNA data in addition to sequencer data such as copy number variation data. The HER2 status prediction computing device 102 may provide both sequencer data (e.g., RNA Seq data) and copy number variation data to a machine learning model trained with a molecular signature. The HER2 status prediction computing device 102 may generate a prediction regarding the sample corresponding to the RNA Seq and copy number variation data.
[0075] In some aspects, database 110 may include additional clinical and molecular patient data. For example, database 110 may include qualitative insights regarding a patient's history, symptoms, and clinical inferences underlying treatment decisions, providing a narrative context to the quantitative data also stored within the database. Database 110 may include laboratory and test results, from basic blood tests, imaging, imaging analysis, to more complex genetic screening. Database 110 may store detailed diagnoses, as well as DNA and RNA sequencing data. In some aspects, database 110 may include methylation assay results, or other information related to epigenetic modifications or disease etiology. In some aspects, database 110 may include treatment response data or other follow-up information related to how a patient responds to various treatments. Database 110 may also incorporate results from methods described in the application, such as risk profiles or HER2 status predictions generated through the application's trained multi-stage machine learning architecture. By integrating these results, the database enhances the utility of molecular data in clinical decision-making, enabling healthcare providers to identify patients who may benefit from targeted therapies based on the patient's HER2 status. This ability represents a significant advancement in the field of oncology where HER2 status is an important factor in determining the most appropriate treatment for cancer patients.
[0076] By the time the HER2 status prediction computing device 102 uses the model thus trained, the trained model has already been trained using a training data set as described above, and one or more submodels that are part of the model architecture (e.g., model architecture 200) are each individually trained to perform binary classification using a labeled data set of RNASeq data and copy number variation data. In some aspects, as mentioned, other data such as proteomics data may be used for training and inference.
[0077] The prediction of the HER2 status prediction computing device 102 can be stored, for example, in the memory 124 or the database 110. These results can be provided directly to other elements of the environment 100, for example, via the network 100. These results can also be further processed, for example, to identify / notify a patient via the patient identification module 156 and / or to generate a digital report by the report generation module 158.
[0078] Exemplary computer-implemented digital report FIG. 3A depicts an exemplary digital report 300 generated by elements of the environment 100 in some embodiments. For example, the report 300 can be generated by the report generation module 158 and displayed via the output device 128 (e.g., via a computer monitor, laptop screen, etc.) and / or via the viewer application 180 of the client computing device 104. FIG. 3A generally depicts a digital report that can be provided to a patient (e.g., via email) that includes the patient's HER2 status as predicted by one or more trained machine learning models (e.g., the model architecture 200 of FIG. 2). The report 300 can include the percentage of likelihood that the patient is HER2 low and the percentage of likelihood that the patient is HER2 negative and / or HER2 positive when assayed by IHC / FISH. In some embodiments, the percentage of likelihood of the model calculation is calculated by comparing the patient's RNA and CNV data to positive and negative controls and determining a numerical estimate of the patient's similarity to the positive and negative controls, respectively.
[0079] The report 300 can include one or more diagnoses (e.g., a diagnosis of cancer such as breast cancer). The report 300 can include additional diagnoses such as estrogen receptor status.
[0080] Figure 3B depicts an exemplary digital report 330 of a patient's progesterone receptor status according to some embodiments. Report 330 may include a PR score that reflects scaled results corresponding to the progesterone receptor status. Report 330 may further include the percentage likelihood that a patient is PR negative or positive as determined by one or more trained machine learning models.
[0081] Figure 3C depicts an exemplary digital report of a patient's gene rearrangement and altered splicing analysis from RNA sequencing according to some embodiments. The report includes a diagnosis of the patient's breast cancer and several sections detailing the results and their potential clinical significance.
[0082] At the top of the report, the patient's name (redacted), breast cancer diagnosis, accession number (hidden), date of birth (hidden), gender (female), sample type (tumor specimen: breast), name of the physician (example of entered text, verification provider), and institution (example of entered text, verification institution - RNA priority) are provided. Also displayed are the address of the laboratory where the test was performed (example of entered text, Journey Test Lab 123,123), collection date (September 14, 2023), and tumor percentage (40%).
[0083] Under the "Genomic Variant" section, "Biologically Relevant" with two findings: "EGFR-SEPT14 chromosomal rearrangement" and "EGFR EGFRvIII altered splicing" is reported. In the "Therapeutic Significance" section, it is stated that "No reportable treatment options were found." In the "Gene Expression" section, there is a description regarding the patient's HER2 overexpression status: "This patient has ERBB2 overexpression and may benefit from HER2 IHC testing if not already performed. In non-solid tumor studies, 81% of samples identified as HER2 overexpressors were 3+ by IHC (95% confidence interval 61% - 89%). The actual positive rate varies by cancer type and depends on the prevalence of HER2 positivity (citation)." Also, a note is issued: "This result is not intended to guide treatment decisions. Potentially actionable results need to be confirmed by clinically appropriate validation tests."
[0084] At the bottom of the report, it includes the electronic signature by the pathologist, CLIA number, date (redacted) on which the report was signed / reported, laboratory medical director, laboratory address (Tempus Labs, Inc. · 600 West Chicago Avenue, Ste 510 · Chicago, IL · 60654 · tempus.com · support@tempus.com), identifier ID, and pipeline version (3.5.3).
[0085] Exemplary computer-implemented method Figure 4 depicts an exemplary flowchart of a computer-implemented method 400 for determining a patient's HER2 low status using the patient's molecular data, according to some embodiments. Method 400 can be implemented using the environment 100 of FIG. 1.
[0086] Method 400 can include receiving digital biological data via one or more processors (block 402).
[0087] Method 400 may include processing digital biometric data corresponding to a patient using a trained multi - stage machine learning architecture via one or more processors, and processing may include using a trained HER2 - positive model to process the digital data to determine whether the digital biometric data indicates that the patient's condition is HER2 - positive, using a trained HER2 - low model to process the digital biometric data to identify whether the patient's condition is HER2 - low when the patient is not HER2 - positive, and designating the patient's condition as HER2 - negative when the patient is neither HER2 - positive nor HER2 - low (block 404). Method 400 may include generating a digital HER2 - low status report corresponding to the patient via one or more processors (block 406). Method 400 may include displaying the digital HER2 - low status report via a display device (block 408). In some aspects, the digital biometric data includes RNA data. In some aspects, the digital biometric data includes at least some transcriptome data. In some aspects, at least some of the transcriptome data includes at least some data generated via RNA seq. In some aspects, the digital biometric data includes at least one of DNA data or copy number variant data. In some aspects, receiving the digital biometric data includes receiving the digital biometric data from a next - generation sequencing platform. In some aspects, the trained HER2 - positive model is a random forest model. In some aspects, the trained HER2 - low model is a random forest model. In some aspects, the trained HER2 - positive model is a binary classifier trained with molecular signature data labeled according to (HER2 - positive, non - HER2 - positive) labels. In some aspects, the trained HER2 - low model is a binary classifier trained with molecular signature data labeled according to (HER2 - low, non - HER2 - low) labels.
[0088] In some aspects, method 400 may include generating a prediction regarding the HER2-low, HER2-positive, and / or HER2-negative status of a given sample based on a trained multi-stage machine learning architecture. In some aspects, method 400 may include identifying at least one patient from a population of patients by processing patient data using a trained multi-stage machine learning architecture, and matching the identified patients for treatment using a targeted therapy. In some aspects, the targeted therapy is a HER2-targeted therapy. In some aspects, the targeted therapy is trastuzumab deruxtecan.
[0089] FIG. 5 depicts a computer-implemented method 500 for training a model architecture to determine a patient's HER2-low status using the patient's molecular data. Method 500 may include receiving, via one or more processors, training digital biological data including a plurality of molecular signatures each having a respective label (block 502). Method 500 may include initializing, via one or more processors, a machine learning model having a plurality of hyperparameters in a computer's memory (block 504). Method 500 may include processing, via one or more processors, using the machine learning model, the plurality of molecular signatures in the training digital biological data and the respective labels of each of the plurality of molecular signatures to generate a trained machine learning model (block 506). Method 500 may include storing, via one or more processors, the trained machine learning model in the computer's memory, which may include generating a serialized copy of the machine learning model and writing the serialized copy of the machine learning model to the computer's memory (block 508).
[0090] In some aspects, computer-implemented method 500 includes loading, via one or more processors, a serialized copy of a trained machine learning into a computer memory, and generating an in-memory instantiation of the trained machine learning model by deserializing the serialized copy of the trained machine learning model, and implementing the method according to claim 1 using the in-memory instantiation of the trained machine learning model.
[0091] In some aspects, the training digital biological data includes at least one of RNA Seq data and copy number variant data. In some aspects, method 500 includes training a machine learning model using at least one RNA feature of at least one of the genes "GRB7", "GSDMB", "MIEN1", "ORMDL3", "PGAP3", "PSMD3", "STARD3", or "ERBB2". In some aspects, method 500 includes training a machine learning model using at least one copy number variant feature of at least one of the genes "BRD4", "CDK12", "ERBB2", or "RARA". In some aspects, the copy number variant data includes, for each copy number variant, the respective estimated copy number variant and the respective estimated confidence of the estimated copy number variant. In some aspects, the machine learning model is a random forest model. Processing, using the machine learning model, a plurality of molecular signatures and labels of each of the plurality of molecular signatures in the training digital biological data to generate a trained machine learning model includes training a first binary classifier to predict a HER2 positive state and training a second binary classifier to predict a HER2 low state.
[0092] In some embodiments, method 500 includes generating a prediction regarding the HER2 low, HER2 positive, and / or HER2 negative status of a given sample based on a trained machine learning model. In some embodiments, method 500 includes identifying at least one patient from a population of patients by processing patient data using a trained machine learning model and matching the identified patients for treatment with a targeted therapy. In some embodiments, the targeted therapy is trastuzumab deruxtecan.
[0093] Additional embodiments can be provided by combining the various embodiments described above. All U.S. patents, U.S. patent application publications, U.S. patent applications, foreign patents, foreign patent applications, and non-patent publications referred to in this specification and / or listed in the Application Data Sheet are hereby incorporated by reference in their entirety. Aspects of the embodiments can be modified as needed to employ concepts from various patents, applications, and publications to provide further additional embodiments.
[0094] In light of the above detailed description, these and other modifications can be made to the embodiments. Generally, in the following claims, the terms used should not be construed as limiting the claims to the specific embodiments disclosed in this specification and the claims, but rather the claims should be construed to include all possible embodiments along with the full scope of equivalents to which such claims are entitled. Accordingly, the claims are not limited by the present disclosure.
[0095] Aspects of the techniques described in this disclosure can include any of the following aspects, alone or in combination.
[0096] 1. A computer-implemented method for determining a patient's HER2-low status using the patient's molecular data, comprising: receiving digital biological data via one or more processors; processing the digital biological data corresponding to the patient using a trained multi-stage machine learning architecture via one or more processors, the processing comprising: (i) processing the digital biological data using a trained HER2-positive model to determine whether the digital biological data indicates that the patient's HER2 status is HER2-positive; (ii) processing the digital biological data using a trained HER2-low model to identify whether the patient's HER2 status is HER2-low when the patient's HER2 status is not HER2-positive; (iii) designating the patient's HER2 status as HER2-negative when the patient's HER2 status is neither HER2-positive nor HER2-low; generating a digital HER2-low status report corresponding to the patient via one or more processors; and displaying the digital HER2-low status report via a display device.
[0097] 2. The computer-implemented method according to aspect 1, wherein the digital biological data includes RNA data.
[0098] 3. The computer-implemented method according to aspect 1 or 2, wherein the digital biological data includes at least some transcriptome data.
[0099] 4. The computer-implemented method according to aspect 3, wherein at least some of the transcriptome data includes at least some data generated via RNA seq.
[0100] 5. The computer-implemented method according to any one of aspects 1 to 4, wherein the digital biological data includes at least one of DNA data or copy number variant data.
[0101] 6. The computer-implemented method according to any one of aspects 1 to 5, wherein receiving digital biological data includes receiving digital biological data from a next-generation sequencing platform.
[0102] 7. The computer-implemented method according to any one of aspects 1 to 6, wherein the trained HER2 positive model is a random forest model.
[0103] 8. The computer-implemented method according to any one of aspects 1 to 7, wherein the trained HER2 low model is a random forest model.
[0104] 9. The computer-implemented method according to any one of aspects 1 to 8, wherein the trained HER2 positive model is a binary classifier trained with molecular signature data labeled according to (HER2 positive, non-HER2 positive) labels.
[0105] 10. The computer-implemented method according to any one of aspects 1 to 9, wherein the trained HER2 low model is a binary classifier trained with molecular signature data labeled according to (HER2 low, non-HER2 low) labels.
[0106] 11. The computer-implemented method according to any one of aspects 1 to 10, further comprising generating a prediction regarding the HER2 low, HER2 positive, and / or HER2 negative state of a given sample based on a trained multi-stage machine learning architecture.
[0107] 12. Identifying at least one patient from a population of patients by processing the patient's data using a trained multi-stage machine learning architecture, and further comprising matching the identified patient for treatment using a targeted therapy. The computer-implemented method according to any one of aspects 1 to 11.
[0108] 13. The computer-implemented method according to aspect 12, wherein the targeted therapy is a HER2 targeted therapy.
[0109] 14. The computer-implemented method according to any one of aspects 11 to 13, wherein the targeted therapy is trastuzumab deruxtecan.
[0110] 15. A computer-implemented method for training a model architecture for determining a patient's HER2 low status using the patient's molecular data, the method comprising: receiving, via one or more processors, training digital biological data including a plurality of molecular signatures each having a respective label; initializing, via one or more processors, a machine learning model having a plurality of hyperparameters in a computer's memory; processing, via one or more processors, using the machine learning model, a plurality of molecular signatures and the label of each of the plurality of molecular signatures in the training digital biological data to generate a trained machine learning model; and storing, via one or more processors, the trained machine learning model in the computer's memory, the storing including generating a serialized copy of the machine learning model and writing the serialized copy of the machine learning model to the computer's memory.
[0111] 16. The computer-implemented method according to aspect 15, further comprising: loading, via one or more processors, a serialized copy of the trained machine learning into the computer's memory; generating an in-memory instantiation of the trained machine learning model by deserializing the serialized copy of the trained machine learning model; and implementing the method of aspect 1 using the in-memory instantiation of the trained machine learning model.
[0112] 17. The computer-implemented method according to aspect 15 or 16, wherein the training digital biological data includes at least one of RNASeq data and copy number variant data.
[0113] 18. The computer-implemented method according to embodiment 17, further comprising training a machine learning model using at least one RNA feature of the genes "GRB7", "GSDMB", "MIEN1", "ORMDL3", "PGAP3", "PSMD3", "STARD3", or "ERBB2".
[0114] 19. The computer-implemented method according to embodiment 17 or 18, further comprising training a machine learning model using at least one copy number variant feature of the genes "BRD4", "CDK12", "ERBB2", or "RARA".
[0115] 20. The computer-implemented method according to any one of embodiments 17 to 19, wherein the copy number variant data includes, for each copy number variant, its respective estimated copy number variant and the respective estimated confidence of the estimated copy number variant.
[0116] 21. The computer-implemented method according to any one of embodiments 15 to 20, wherein the machine learning model is a random forest model.
[0117] 22. In order to generate a trained machine learning model, using the machine learning model to process a plurality of molecular signatures and the label of each of the plurality of molecular signatures in the training digital biological data, including training a first binary classifier to predict the HER2 positive state and training a second binary classifier to predict the HER2 low state. The computer-implemented method according to any one of embodiments 15 to 20.
[0118] 23. The computer-implemented method according to any one of embodiments 15 to 22, further comprising generating a prediction regarding the HER2 low, HER2 positive, and / or HER2 negative state of a given sample based on the trained machine learning model.
[0119] 24. Further comprising identifying at least one patient from a population of patients by processing patient data using a trained machine learning model, and matching the identified patients for treatment using a targeted therapy, the computer-implemented method according to any one of aspects 15 to 23.
[0120] 25. The computer-implemented method according to any one of aspects 15 to 24, wherein the targeted therapy is trastuzumab deruxtecan.
[0121] 26. A computing system comprising one or more processors and one or more memories storing computer-readable instructions, the computer-readable instructions, when executed, causing the computing system to receive digital biometric data, and process the digital biometric data corresponding to a patient using a trained multi-stage machine learning architecture, the processing including: (i) processing the digital data using a trained HER2 positive model to determine whether the digital biometric data indicates that the patient's HER2 status is HER2 positive; (ii) processing the digital biometric data using a trained HER2 low model to identify whether the patient's HER2 status is HER2 low when the patient is not HER2 positive; and (iii) designating the patient's HER2 status as HER2 negative when the patient is neither HER2 positive nor HER2 low, and generating, via one or more processors, a digital HER2 low status report corresponding to the patient, and displaying the digital HER2 low status report via a display device.
[0122] 27. A computer-readable medium storing computer-executable instructions that, when executed, cause a computer to receive digital biological data and process the digital biological data corresponding to a patient using a trained multi-stage machine learning architecture, the processing including: (i) processing the digital data using a trained HER2 positive model to determine whether the digital biological data indicates that the patient's HER2 status is HER2 positive; (ii) processing the digital biological data using a trained HER2 low model to identify whether the patient's HER2 status is HER2 low when the patient is not HER2 positive; and (iii) designating the patient's HER2 status as HER2 negative when the patient is neither HER2 positive nor HER2 low, and generating a digital HER2 low status report corresponding to the patient via one or more processors and causing the digital HER2 low status report to be displayed via a display device.
[0123] 28. A computing system comprising one or more processors and one or more memories storing computer-readable instructions that, when executed, cause the computing system to receive training digital biological data including a plurality of molecular signatures each having a respective label, initialize a machine learning model having a plurality of hyperparameters in the computer's memory, process the plurality of molecular signatures and the respective labels of the plurality of molecular signatures in the training digital biological data using the machine learning model to generate a trained machine learning model, and store the trained machine learning model in the computer's memory, the storing including generating a serialized copy of the machine learning model and writing the serialized copy of the machine learning model to the computer's memory.
[0124] 29. A computer-readable medium storing computer-executable instructions that, when executed, cause a computer to receive training digital biological data including a plurality of molecular signatures each having a respective label, initialize a machine learning model having a plurality of hyperparameters in a memory of the computer, process the plurality of molecular signatures and the label of each of the plurality of molecular signatures in the training digital biological data using the machine learning model to generate a trained machine learning model, and store the trained machine learning model in the memory of the computer, the storing including generating a serialized copy of the machine learning model and writing the serialized copy of the machine learning model to the memory of the computer.
[0125] Additional Considerations The computer-readable medium may include executable computer-readable code stored on a computer (e.g., including a processor and a GPU) for programming the computer with the techniques of this specification. Examples of such computer-readable storage media include hard disks, CD-ROMs, digital versatile disks (DVDs), optical storage devices, magnetic storage devices, ROM (read-only memory), PROM (programmable read-only memory), EPROM (erasable programmable read-only memory), EEPROM (electrically erasable programmable read-only memory), and flash memory. More generally, the processing unit of computing device 1300 may represent a CPU-type processing unit, a GPU-type processing unit, a TPU-type processing unit, a field programmable gate array (FPGA), another class of digital signal processor (DSP), or other hardware logic components drivable by a CPU.
[0126] A system for executing the methods described herein may include a computing device, and more specifically, may be implemented on one or more processing units, such as a central processing unit (CPU), and / or one or more graphics processing units (GPUs) including a cluster of CPUs and / or GPUs. The features and functions described may be stored on one or more non-transitory computer-readable media of the computing device and then implemented. The computer-readable media may include, for example, an operating system and software modules, or “engines,” that implement the methods described herein. These engines may be stored as a set of non-transitory computer-executable instructions. The computing device may be a distributed computing system such as Amazon Web Services, Google Cloud Platform, Microsoft Azure, or other public, private, and / or hybrid cloud computing solutions.
[0127] The computing device includes a network interface communicatively coupled to a network for communicating to and / or from a portable personal computer, smartphone, electronic document, tablet, and / or desktop personal computer, or other computing device. The computing device further includes an I / O interface connected to devices such as a digital display, user input device, and the like.
[0128] The functions of the engine can be implemented in distributed computing devices interconnected with each other via a communication link, etc. In other embodiments, the functions of the system can be distributed across any number of devices, including the portable personal computer, smartphone, electronic document, tablet, and desktop personal computer devices shown. The computing devices can be communicatively coupled to a network and another network. The network can be a public network such as the Internet, a private network such as a network of a research institution or a company, or any combination thereof. The network includes local area network (LAN), wide area network (WAN), cellular, satellite, or other network infrastructure, whether wireless or wired. The network can utilize communication protocols including packet-based and / or datagram-based protocols such as Internet Protocol (IP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), or other types of protocols. Further, the network can include a plurality of devices such as switches, routers, gateways, access points (such as wireless access points as shown), firewalls, base stations, repeaters, backbone devices, etc., that facilitate network communication and / or form the hardware infrastructure of the network.
[0129] A computer-readable medium may include executable computer-readable code stored on a computer for programming a computer with the techniques of this specification (e.g., including a processor and a GPU). Examples of such computer-readable storage media include hard disks, CD-ROMs, digital versatile disks (DVDs), optical storage devices, magnetic storage devices, ROM (read-only memory), PROM (programmable read-only memory), EPROM (erasable programmable read-only memory), EEPROM (electrically erasable programmable read-only memory), and flash memory. More generally, the processing unit of a computing device may represent a CPU-type processing unit, a GPU-type processing unit, a field programmable gate array (FPGA), another class of digital signal processor (DSP), or other hardware logic components that can be driven by a CPU.
[0130] Throughout this specification, multiple instances may implement components, operations, or structures described as a single instance. Although individual operations of one or more methods are illustrated and described as separate operations, one or more of the individual operations may be performed simultaneously and need not be performed in the order illustrated. Structures and functions presented as separate components within an exemplary configuration may be implemented as a combined structure or component. Similarly, structures and functions presented as a single component may be implemented as separate components or a plurality of components.
[0131] Furthermore, certain aspects are described herein as including logic or some routines, sub-routines, applications, or instructions. These can constitute either software (e.g., code embodied on a machine-readable medium or in a transmission signal) or hardware. In hardware, routines and the like are tangible units capable of performing certain operations and can be configured or arranged in a particular manner. In an exemplary aspect, one or more computer systems (e.g., a stand-alone, client, or server computer system), or one or more hardware modules of a computer system (e.g., a processor or group of processors), can be configured as hardware modules that operate to perform the particular operations described herein by software (e.g., an application or a part of an application).
[0132] In various aspects, the hardware modules can be implemented mechanically or electronically. For example, a hardware module can comprise dedicated circuitry or logic (e.g., a special-purpose processor such as a microcontroller, a field programmable gate array (FPGA), or an application specific integrated circuit (ASIC)) that is permanently configured to perform certain operations. A hardware module can also comprise programmable logic or circuitry (e.g., that included within a general-purpose processor or other programmable processor) that is temporarily configured by software to perform certain operations. It will be appreciated that whether to implement a hardware module mechanically, with dedicated and permanently configured circuitry, or with temporarily configured circuitry (e.g., configured by software) can be determined considering cost and time.
[0133] Accordingly, the term "hardware module" is to be understood as encompassing a tangible entity, which is physically constructed or permanently configured (e.g., embedded in hardware) or temporarily configured (e.g., programmed) to operate in a particular manner or to perform certain operations described herein. Considering the case where a hardware module is temporarily configured (e.g., programmed), each of the hardware modules need not be configured or instantiated at any given point in time. For example, if a hardware module includes a general-purpose processor configured using software, the general-purpose processor can be configured as different hardware modules at different points in time. Thus, software can configure the processor, for example, to constitute a particular hardware module at one point in time and a different hardware module at another point in time.
[0134] A hardware module can provide information to and receive information from other hardware modules. Accordingly, the described hardware modules can be considered to be communicatively coupled. If multiple such hardware modules exist simultaneously, communication can be achieved via signal transmission connecting the hardware modules (e.g., via appropriate circuitry and buses). In the case where multiple hardware modules are configured or instantiated at different times, communication between such hardware modules can be achieved, for example, via storage and retrieval of information in a memory structure accessed by the multiple hardware modules. For example, a hardware module can perform an operation and store the output of the operation in a memory device to which the hardware module is communicatively coupled. Subsequently, a further hardware module can later access the memory device to retrieve and process the stored output. A hardware module can also initiate communication with an input or output device and operate on resources (e.g., collect information).
[0135] The various operations of the exemplary methods described herein can be performed, at least in part, by one or more processors temporarily configured (e.g., by software) or permanently configured to perform the associated operations. Whether temporarily or permanently configured, such processors can constitute processor-implemented modules that operate to perform one or more operations or functions. Modules referred to herein can, in some exemplary aspects, comprise processor-implemented modules.
[0136] Similarly, the methods or routines described herein can be at least in part processor-implemented. For example, at least some of the operations of a method can be performed by one or more processors or processor-implemented hardware modules. Certain performance of operations can exist not only within a single machine but also be distributed among one or more processors deployed across several machines. In some exemplary aspects, one or more processors can be located in a single location (e.g., within a home environment, within a workplace environment, or as a server farm), but in other aspects, the processors can be distributed across a number of locations.
[0137] Certain performance of operations can exist not only within a single machine but also be distributed among one or more processors deployed across several machines. In some exemplary aspects, one or more processors or processor-implemented modules can be located in a single location (e.g., within a home environment, within a workplace environment, or as a server farm). In other exemplary aspects, one or more processors or processor-implemented modules can be distributed across a number of locations.
[0138] Unless otherwise specified, the discussions in this specification using terms such as "processing", "computing", "calculating", "determining", "presenting", "displaying", etc. may refer to operations or processes of a machine (e.g., a computer) that manipulates or transforms data represented as a physical (e.g., electronic, magnetic, or optical) quantity within one or more memories (e.g., volatile memory, non-volatile memory, or a combination thereof), registers, or other mechanical components that receive, store, transmit, or display information.
[0139] As used herein, any reference to "one aspect" or "an aspect" means that a particular element, feature, structure, or characteristic described in conjunction with that aspect is included in at least one aspect. The appearances of the phrase "in one aspect" in various places in this specification are not necessarily all referring to the same aspect.
[0140] Some aspects may be described using the expressions "coupled" and "connected" along with their derivatives. For example, some aspects may be described using the term "coupled" to indicate that two or more elements are in direct physical or electrical contact. However, the term "coupled" may also mean that two or more elements are not in direct contact with each other but still cooperate or interact with each other. The aspects are not limited to this context.
[0141] As used herein, the terms "comprises," "comprising," "includes," "including," "has," "having," or any other variation thereof are intended to cover a non-exclusive inclusion. For example, a process, method, article, or apparatus that comprises a list of elements is not necessarily limited to only those elements, but may include other elements not expressly listed or inherent to such process, method, article, or apparatus. Further, unless expressly stated to the contrary, "or" as used herein means "and / or" and not "either / or." For example, condition A or B is satisfied by any one of the following: A is true (or present) and B is false (or not present), A is false (or not present) and B is true (or present), and both A and B are true (or present).
[0142] In addition, the use of "a" or "an" is employed to describe elements and components of aspects of this specification. This is done merely for convenience and to give a general sense of the description. This description should be read to include one or at least one, and singular also includes plural unless it is obvious that the contrary is meant.
[0143] This detailed description should be construed as illustrative only and not as exhaustive, since it is impractical, if not impossible, to describe all possible embodiments. Many alternative embodiments can be implemented using any of the techniques developed after the filing date of this technology or this patent application.
Claims
1. 1. A computer-implemented method for determining HER2 low status in a patient using molecular data of said patient, comprising: receiving, via one or more processors, digital biometric data; processing the digital biometric data corresponding to the patient using a trained multi-stage machine learning architecture via one or more processors, wherein said processing includes: (i) processing the digital biometric data using a trained HER2-positive model to determine whether the digital biometric data is indicative of the patient's HER2 status being HER2-positive; (ii) processing the digital biometric data using a trained HER2 low model to identify whether the HER2 status of the patient is HER2 low when the HER2 status of the patient is not HER2 positive; (iii) when the HER2 status of the patient is not HER2 positive or HER2 low, designating the HER2 status of the patient as HER2 negative; generating, via one or more processors, a digital HER2 low status report corresponding to said patient; and displaying the digital HER2 low status report via a display device.
2. The computer-implemented method of claim 1 , wherein the digital biometric data comprises RNA data.
3. The computer-implemented method of claim 1 , wherein the digital biometric data includes at least some transcriptomic data.
4. 4. The computer-implemented method of claim 3, wherein the at least some of the transcriptomic data comprises at least some data generated via RNA seq.
5. The computer-implemented method of claim 1 , wherein the digital biometric data comprises at least one of DNA data or copy number variant data.
6. 10. The computer-implemented method of claim 1, wherein receiving the digital biometric data comprises receiving the digital biometric data from a next-generation sequencing platform.
7. 2. The computer-implemented method of claim 1, wherein the trained HER2-positive model is a random forest model.
8. 2. The computer-implemented method of claim 1 , wherein the trained HER2 low model is a random forest model.
9. 2. The computer-implemented method of claim 1, wherein the trained HER2-positivity model is a binary classifier trained on molecular signature data labeled according to (HER2-positive, non-HER2-positive) labels.
10. 2. The computer-implemented method of claim 1, wherein the trained HER2 low model is a binary classifier trained on molecular signature data labeled according to (HER2 low, non-HER2 low) labels.
11. 2. The computer-implemented method of claim 1, further comprising generating a prediction regarding the HER2-low, HER2-positive, and / or HER2-negative status of a given sample based on the trained multi-stage machine learning architecture.
12. identifying at least one patient from a population of patients by processing said data of said patient using a trained multi-stage machine learning architecture; 10. The computer-implemented method of claim 1, further comprising matching the identified patient for treatment with a targeted therapy.
13. 13. The computer-implemented method of claim 12, wherein the targeted therapy is a HER2 targeted therapy.
14. 14. The computer-implemented method of claim 13, wherein the targeted therapy is trastuzumab derquistecan.
15. 1. A computing system comprising: one or more processors; and one or more memories storing computer readable instructions that, when executed, cause the computing system to: Receiving digital biometric data; processing the digital biometric data corresponding to a patient using a trained multi-stage machine learning architecture, said processing including: (i) processing the digital biometric data using a trained HER2-positive model to determine whether the digital biometric data is indicative of the patient's HER2 status being HER2-positive; (ii) processing the digital biometric data using a trained HER2 low model to identify whether the HER2 status of the patient is HER2 low when the patient is not HER2 positive; (iii) when the patient is not HER2 positive or HER2 low, designating the HER2 status of the patient as HER2 negative; generating, via one or more processors, a digital HER2 low status report corresponding to said patient; and displaying said digital HER2 low status report via a display device.
16. 16. The computing system of claim 15, wherein the digital biological data includes one or both of: (i) RNA data; and (ii) at least some transcriptome data.
17. 16. The computing system of claim 15, wherein the at least some of the transcriptomic data comprises at least some data generated via RNA seq.
18. A computer-readable medium having stored thereon computer-executable instructions that, when executed, cause a computer to: Receiving digital biometric data; processing the digital biometric data corresponding to a patient using a trained multi-stage machine learning architecture, said processing including: (i) processing the digital biometric data using a trained HER2-positive model to determine whether the digital biometric data is indicative of the patient's HER2 status being HER2-positive; (ii) processing the digital biometric data using a trained HER2 low model to identify whether the patient's HER2 status is HER2 low when the patient is not HER2 positive; (ii) when the patient is not HER2 positive or HER2 low, designating the HER2 status of the patient as HER2 negative; generating, via one or more processors, a digital HER2 low status report corresponding to said patient; and displaying said digital HER2 low status report via a display device.
19. 20. The computer-readable medium of claim 18, wherein the digital biological data includes one or both of: (i) RNA data; and (ii) at least some transcriptome data.
20. 20. The computer-readable medium of claim 18, wherein the at least some of the transcriptomic data comprises at least some data generated via RNA seq.