Identifying mutation signatures using machine learning

A machine learning model predicts mutational signatures from a small number of mutations using embeddings and natural language processing, addressing the limitations of current methods by enabling accurate clinical predictions from limited sequencing data.

JP2025535869APending Publication Date: 2025-10-30HADASIT MEDICAL RESEARCH SERVICES & DEVELOPMENT LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025515835
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-15
Filing Date
2023-09-14
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Current methods for identifying mutational signatures in cancer are limited by the need for whole-genome sequencing, which is not feasible in clinical settings due to the high cost and low mutation coverage of targeted gene panels, making it difficult to accurately predict mutational signatures from a limited number of mutations.

Method used

A machine learning model is trained to identify dominant mutation signatures using a cohort of DNA sequencing samples, learning embeddings for mutations and signatures, allowing prediction from a small number of mutations, even as few as one, by utilizing natural language processing techniques to associate mutations with their signatures and tumor types.

Benefits of technology

The model accurately predicts mutational signatures from limited sequencing data, enabling clinical applications such as personalized cancer treatment and survival prediction, even with as few as four mutations, and demonstrates robust performance across various datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025535869000001_ABST
    Figure 2025535869000001_ABST
Patent Text Reader

Abstract

The goal is to improve our ability to identify mutational signatures. [Solution] A computer-implemented method includes receiving as input a plurality of DNA sequencing samples corresponding to a cohort of subjects having various tumor types, annotating each of the DNA sequencing samples to indicate (a) the tumor type associated with the DNA sequencing sample, (b) one or more mutations present in the DNA sequencing sample, and (c) at least one dominant mutation signature associated with the DNA sequencing sample, training a machine learning model in a training phase using a training dataset including (i) the DNA sequencing samples and (ii) labels indicating the annotations, and applying the trained machine learning model to target DNA sequencing samples from target subjects to predict dominant mutation signatures associated with the target DNA sequencing samples in an inference phase.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Application No. 63 / 406,909, filed September 15, 2022, entitled "MACHINE LEARNING IDENTIFICATION OF MUTATIONAL SIGNATURES," the contents of which are incorporated herein by reference in their entirety.

[0002]

[0002] The present invention relates to the field of machine learning. [Background technology]

[0003] Somatic mutations accumulate through multiple mutational processes, forming patterns called "mutational signatures." In cancer, mutational signatures reflect causal processes. Some somatic mutations, such as SBS1 and SBS5, are endogenous, ubiquitous, and age-related. Other mutations reflect specific exogenous or endogenous processes, such as UV-associated signatures and APOBEC-associated signatures.

[0004] Mutational signatures can provide important insights into prognosis and treatment. For example, in one clinical application, the APOBEC signature was identified as a strong and specific predictor of resistance to EGFR-TKIs. Similarly, patients with APOBEC or homologous recombination deficiency signatures may benefit from ATR inhibitors or PARP inhibitors. In another example, immune checkpoint inhibitors may be beneficial for tumors with UV, Tobacco, APOBEC, POLE, or MMR signatures.

[0005]

[0005] Deciphering mutational signatures in cancer provides insight into the biological mechanisms involved in carcinogenesis and normal somatic mutagenesis. Mutational signatures have potential applications in cancer treatment and prevention. Advances in the field of tumor genomics have enabled the development and use of molecularly targeted therapies. More recently, mutational signatures have been shown to be associated with therapeutic response and may serve as biomarkers to predict response to antitumor drugs.

[0006] Despite these promising advances, the use of mutational signatures in clinical settings is limited. This is primarily because mutational processes are context-dependent rather than gene-specific, so signature detection requires analyzing the mutational landscape across the entire genome. Furthermore, identifying a mutational signature requires a sufficient number of mutations. Therefore, whole-genome sequencing (WGS) is typically considered the gold standard for the signature framework, although whole-exome sequencing (WES) may be sufficient in some cases. However, currently, WGS and WES are primarily used for research purposes.

[0007] Conversely, targeted gene panels are routinely performed on patients diagnosed with cancer, with millions of targeted gene panels administered annually. However, targeted gene panels only cover a maximum of 2 Mb (<0.1%) of the genome, capturing only a small fraction of the mutational landscape and resulting in a low number of mutations per sample.

[0008]

[0008] Thus, the ability to identify mutational signatures in targeted gene panels remains an unmet need.

[0009]

[0009] The foregoing examples of the related art and their associated limitations are intended to be illustrative and not exhaustive. Further limitations of the related art will become apparent to those skilled in the art upon reading the specification and studying the drawings. Summary of the Invention

[0010]

[0010] The following embodiments and aspects thereof are described and illustrated with reference to systems, means and methods that are exemplary and illustrative, not limiting in scope.

[0011]

[0011] In one embodiment, a computer-implemented method is provided that includes receiving as input a plurality of DNA sequencing samples corresponding to a cohort of subjects having various tumor types; annotating each of the DNA sequencing samples to indicate (a) the tumor type associated with the DNA sequencing sample, (b) one or more mutations present in the DNA sequencing sample, and (c) at least one dominant mutation signature associated with the DNA sequencing sample; during a training phase, training a machine learning model using a training dataset including (i) the DNA sequencing samples and (ii) labels indicating the annotations; and during an inference phase, applying the trained machine learning model to target DNA sequencing samples from target subjects to predict dominant mutation signatures associated with the target DNA sequencing samples.

[0012]

[0012] In one embodiment, there is also provided a system including at least one hardware processor and a non-transitory computer-readable storage medium having program instructions stored thereon, wherein the program instructions executable by the at least one hardware processor receive as input a plurality of DNA sequencing samples corresponding to a cohort of subjects having different tumor types, annotate each of the DNA sequencing samples indicating (a) the tumor type associated with the DNA sequencing sample, (b) one or more mutations appearing in the DNA sequencing sample, and (c) at least one dominant mutation signature associated with the DNA sequencing sample, during a training phase, training a machine learning model using a training dataset including (i) the DNA sequencing samples and (ii) labels indicating the annotations, and during an inference phase, applying the trained machine learning model to target DNA sequencing samples from target subjects to predict dominant mutation signatures associated with the target DNA sequencing samples.

[0013]

[0013] In one embodiment, there is further provided a computer program product including a non-transitory computer-readable storage medium having program instructions embodied therein, the program instructions being executable by at least one hardware processor, which receives as input a plurality of DNA sequencing samples corresponding to a cohort of subjects having different tumor types, annotates each of the DNA sequencing samples indicating (a) the tumor type associated with the DNA sequencing sample, (b) one or more mutations appearing in the DNA sequencing sample, and (c) at least one dominant mutation signature associated with the DNA sequencing sample, during a training phase, training a machine learning model using a training dataset including (i) the DNA sequencing samples and (ii) labels indicating the annotations, and during an inference phase, applying the trained machine learning model to target DNA sequencing samples from target subjects to predict dominant mutation signatures associated with the target DNA sequencing samples.

[0014]

[0014] In some embodiments, the training phase includes pre-learning an embedding in an n-dimensional vector space by a machine learning model for each of the dominant mutation signatures and mutations represented in the DNA sequencing samples.

[0015]

[0015] In some embodiments, the inference step includes (i) creating a vector embedding in an n-dimensional vector space for each mutation represented in the target DNA sequencing sample; (ii) calculating a similarity score between (a) the average mutation embedding vector of all vector embeddings created for each mutation represented in the target DNA sequencing sample and (b) each of the dominant mutation signature embeddings pre-trained in the training step; and (iii) selecting the pre-trained dominant mutation signature embedding showing the highest similarity score as the predicted dominant mutation signature associated with the target DNA sequencing sample.

[0016] In some embodiments, the similarity score is based on a cosine similarity calculation.

[0017]

[0017] In some embodiments, the targeted DNA sequencing sample contains 1 to 15 mutations.

[0018]

[0018] In some embodiments, the annotation further indicates, for at least a portion of the DNA sequencing sample, at least one of the following annotation categories: tumor site, metastatic site, structural mutation, smoking history of the corresponding subject, vital status of the corresponding subject, vital status of the corresponding subject, gender of the corresponding subject, or age of the corresponding subject.

[0019] In some embodiments, annotations for mutations represented in DNA sequencing samples are represented by 5-mer sequences.

[0020] In some embodiments, for at least some of the DNA sequencing samples, the annotation represents a combination of two dominant mutation signatures.

[0021]

[0021] In some embodiments, the training step further includes pre-training a joint embedding for each combination of two dominant mutation signatures.

[0022]

[0022] In some embodiments, the inference step includes (i) creating a vector embedding in an n-dimensional vector space for each mutation represented in the target DNA sequencing sample; (ii) calculating a similarity score between (a) the average mutation embedding vector of all vector embeddings created for each mutation represented in the target DNA sequencing sample and each of (b) the pre-trained dominant mutation signature embedding and the combined embedding; and (iii) selecting the (a) pre-trained dominant mutation signature embedding or (b) combined embedding that exhibits the highest similarity score as the predicted dominant mutation signature associated with the target DNA sequencing sample.

[0023]

[0023] In one embodiment, there is further provided a method of treating cancer in a patient in need of treatment, comprising the steps of obtaining a DNA sequencing sample from the patient, applying a computer-implemented method described in any one of claims 1 to 10 to predict a dominant mutation signature associated with the DNA sequencing sample, and administering cancer treatment to the patient based at least in part on the predicted dominant mutation signature.

[0024]

[0024] In one embodiment, there is further provided a method for predicting survival of a cancer patient, comprising the steps of obtaining a DNA sequencing sample from the patient, applying a computer-implemented method described in any one of claims 1 to 10 to predict a dominant mutation signature associated with the DNA sequencing sample, and predicting survival of the patient based at least in part on the predicted dominant mutation signature.

[0025]

[0025] In one embodiment, there is further provided a method for predicting immunotherapy response in a cancer patient, comprising the steps of obtaining a DNA sequencing sample from the patient, applying a computer-implemented method described in any one of claims 1 to 10 to predict a dominant mutation signature associated with the DNA sequencing sample, and predicting the patient's immunotherapy response based at least in part on the predicted dominant mutation signature.

[0026]

[0026] In one embodiment, there is further provided a method for treating cancer in a patient in need of treatment, comprising the steps of obtaining a DNA sequencing sample from the patient; applying to the DNA sequencing sample a machine learning model trained to predict a dominant mutation signature associated with the DNA sequencing sample; and administering cancer treatment to the patient based at least in part on the dominant mutation signature in the DNA sequencing sample predicted by the application.

[0027]

[0027] In one embodiment, there is further provided a method for predicting survival of a cancer patient, comprising the steps of obtaining a DNA sequencing sample from the patient, applying to the DNA sequencing sample a machine learning model trained to predict dominant mutation signatures associated with the DNA sequencing sample, and predicting survival of the patient based on at least a portion of the dominant mutation signatures in the DNA sequencing sample predicted by the application.

[0028]

[0028] In one embodiment, there is further provided a method for predicting immunotherapy response in a cancer patient, the method comprising the steps of obtaining a DNA sequencing sample from the patient; applying to the DNA sequencing sample a machine learning model trained to predict dominant mutation signatures associated with the DNA sequencing sample; and predicting the patient's immunotherapy response based on at least a portion of the dominant mutation signatures in the DNA sequencing sample predicted by the application.

[0029]

[0029] In addition to the exemplary aspects and embodiments described above, further aspects and embodiments will become apparent by reference to the drawings and by study of the detailed description below. [Brief explanation of the drawings]

[0030] [Figure 1A]

[0030] An exemplary training workflow for the present machine learning model according to some embodiments of the present invention is shown. [Figure 1B]

[0031] 1B shows a visualization of the signature embeddings produced by the training process illustrated in FIG. 1A according to some embodiments of the present invention. [Figure 1C]

[0032] 1 illustrates an exemplary inference process for a trained machine learning model of the present disclosure, in accordance with some embodiments of the present invention. [Figure 2]

[0033] FIG. 1 is a block diagram of an exemplary system for training a machine learning model configured to identify specific mutational signatures based on a limited number of input mutations, according to some embodiments of the present invention. [Figure 3]

[0034] 1 illustrates functional steps in a method for training a machine learning model configured to identify dominant mutation signatures in DNA sequencing samples obtained from a subject, according to some embodiments of the present invention. [Figures 4A-4D]

[0035] 1 shows experimental results of the present machine learning model predicting mutational signatures using down-sampled target samples, according to some embodiments of the present invention. [Figure 5A-5B]

[0036] 1 shows experimental results of the present machine learning model predicting mutation signatures in PCAWG WGS samples, according to some embodiments of the present invention. [Figures 6A-6H]

[0037] 1 shows experimental results of the present machine learning model predicting mutational signatures in the MSK-IMPACT cohort, according to some embodiments of the present invention. [Figures 7A-7F]

[0038] 1 shows experimental clinical results of the present machine learning model predicting mutational signatures in the MSK-ICI and MSK-MET cohorts, according to some embodiments of the present invention. [Figure 8A-8B]

[0039] 1 shows experimental clinical results of the present machine learning model predicting mutational signatures in the MSK-ICI and MSK-MET cohorts, according to some embodiments of the present invention. [Figure 9A-9B]

[0040] 7A-7F and 8A-8B show comparisons of p-values ​​and hazard ratio values ​​between results of multivariate survival analyses (as shown in FIGS. 7A-7F and 8A-8B), according to some embodiments of the present invention. [Figures 10A-10E]

[0041] 1 shows survival rates of NSCLC patients according to some embodiments of the present invention. [Figures 11A-11B]

[0042] 1 illustrates the interpretability of mutation and UV signature embedding representations according to some embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0031]

[0043] Techniques embodied in a system, a computer-implemented method, and a computer program product are disclosed for a machine learning model trained to identify dominant mutation signatures in DNA sequencing samples obtained from a subject. In some embodiments, the machine learning model of the present disclosure may be trained to identify dominant mutation signatures in DNA sequencing samples obtained from a subject, and the DNA sequencing samples may consist of only a limited amount of sequencing data, such as sequencing samples obtained using a targeted gene panel.

[0032]

[0044] As used herein, the term "mutation signature" or "signature" refers to the mutation pattern that is characteristic of a mutagenesis process as a whole.

[0033]

[0045] In some embodiments, the machine learning model, during a training phase, learns embeddings representative of mutations and mutational signatures associated with various tumor types, as well as the contextual relationships between the mutations and their signatures. In some embodiments, the trained machine learning model, during an inference phase, may be applied to targeted sequencing samples obtained from a subject to predict dominant signatures associated with the target samples. In some embodiments, the predictions are based, at least in part, on the individual mutations present in the target samples. In some embodiments, the target samples may consist of only a limited amount of sequencing data, such as a limited number of mutations.

[0034]

[0046] In some embodiments, the machine learning model learns the context of mutations and signatures based on a training dataset that includes multiple tumor sequencing samples from a cohort of subjects, where each sample is annotated to indicate at least some of the following: Tumor type The class of mutation represented in the sample (for each mutation, considering the pentanucleotide context, e.g., the mutation event and the two bases at the 5' and 3' ends, there are a total of 1536 possible options) Dominant signatures associated with the sample, e.g., APOBEC (SBS2, SBS13), UV (SBS7a-d and SBS38), Tobacco (SBS4), HRD (SBS3), MMR (SBS6, 14, 15, 20, 21, 26, and 44), POLE (SBS10a-b), Clock_SBS1, and Clock_SBS5

[0035]

[0047] In some embodiments, the machine learning model incorporates natural language processing (NLP) machine learning techniques based on the insight that DNA sequencing samples can be viewed as documents containing individual mutations that can be likened to words or tokens. Meanwhile, the dominant mutation signatures active within a sample and their associated cancer types can be likened to "tags" assigned to the entire document. Thus, during training, the machine learning model aims to create a numerical representation (or embedding) for each DNA training sample based on the individual mutations and the assigned "tags" of the dominant signatures and associated tumor types. This maximizes the similarity between related features (e.g., UV-damage mutation embedding and UV signature embedding) while simultaneously minimizing the similarity between unrelated features (e.g., UV signature embedding and Tobacco signature embedding).

[0036]

[0048] In some embodiments, a trained machine learning model of the present disclosure may be inferred on target samples to create target vector embeddings, which may then be compared to the signature embeddings created during training to predict whether they appear in the learned embeddings.

[0037]

[0049] Our experiments showed that the predictions of this machine learning model were robust in a variety of settings, including WGS samples, downsampled WES (downsampling to 20% of the mutation burden or 1-15 mutations), and when inferred on targeted gene panels with high tumor mutation burden.

[0038]

[0050] The potential advantage of the present invention is that it can accurately predict the association between a specific mutation and a mutation signature based on a sample containing a relatively small number of mutations, for example, four mutations, or even one mutation depending on the signature.Therefore, the present invention may have important clinical significance in situations where only a limited amount of DNA sequencing data is available, such as in the case of general clinical tests such as targeted gene panels.The present invention provides an easy-to-use method for detecting dominant signatures with high accuracy in clinical testing, which may lead to better personalized medicine for cancer patients.

[0039]

[0051] Figure 1A shows an example training workflow for our machine learning model. The model is trained using a training dataset (left) containing tumor sequencing samples annotated with the underlying mutations, dominant signatures, and associated tumor types. The model training process (middle) aims to reduce the loss while maximizing the similarity between relevant features and minimizing the similarity between unrelated features. The result is a set of embeddings of various signatures and associated mutations in a Cartesian coordinate system (right).

[0040]

[0052] Figure 1B shows a visualization of the signature embeddings produced by the training process illustrated in Figure 1A, which have been reduced to two dimensions by applying the t-distributed stochastic neighbor embedding (t-SNE projection) algorithm. As shown, the signature embeddings are clearly clustered in the visualized embedding space.

[0041]

[0053] 1C illustrates an exemplary inference process for a trained machine learning model of the present disclosure. An embedding representation is created for a target sample containing, in this example, four mutations. The cosine similarity between the created target profile and each signature embedding created during training is then calculated. The embedding with the highest cosine similarity to the target profile embedding (where the cosine is greater than a predetermined threshold, e.g., 0.6) is predicted to be the dominant signature active in the target sample.

[0042]

[0054] FIG. 2 is a block diagram of an exemplary system 200 for training a machine learning model configured to identify specific mutational signatures based on a limited number of input mutations, according to some embodiments of the present invention.

[0043]

[0055] In some embodiments, system 200 may include a hardware processor 202 and random access memory (RAM) 204, and / or one or more non-transitory computer-readable storage devices 206. In some embodiments, system 200 may store software instructions or components in storage device 206 configured to operate a processing unit, such as hardware processor 202 (also referred to as a “hardware processor,” “CPU,” “quantum computer processor,” or simply “processor”). In some embodiments, the software components may include an operating system including various software components and / or drivers used to control and manage general system tasks (e.g., memory management, storage device control, power management, etc.) and facilitate communication between the various hardware and software components. The components of system 200 may be co-located or distributed, or the system may be configured to run as one or more cloud computing “instances,” “containers,” “virtual machines,” or another type of encapsulated software application, as known in the art.

[0044]

[0056] The storage device 206 may store program instructions and / or components configured to operate the hardware processor 202. The program instructions may include one or more software modules, such as a machine learning module 206a, an embedding module 206b, and a signature prediction module 206c.

[0045]

[0057] The machine learning module 206a may include any one or more suitable neural network architectures (i.e., including one or more neural network layers) and may be implemented using any suitable optimization algorithm. In some embodiments, the machine learning module 206a may be configured to train a machine learning model of the present disclosure to identify dominant mutation signatures in a DNA sequencing target sample 220 obtained from a subject.

[0046]

[0058] In some embodiments, the machine learning module 206a may utilize any one or more natural language processing (NLP) and / or clustering algorithms, such as genetic clustering algorithms, Bayesian clustering, C-means, K-means, fuzzy logic, etc.

[0047]

[0059] In some embodiments, the machine learning model trained by the machine learning module 206a may be configured to perform clustering operations on the DNA sequencing sample input data to pre-learn embeddings for individual mutations based on their associated mutational signatures. In some embodiments, the machine learning model trained by the machine learning module 206a may incorporate an embedding module 206b configured to generate a token embedding for each mutation identified in the input sequencing sample, mapping each mutation to an n-dimensional vector space based on its context. Thus, mutations with similar contexts will appear in approximately the same location in the vector space. Each mutation in the input sequencing sample is assigned a numerical vector that represents the mutation in the embedding space, which itself captures the relationship between various mutations represented in the input data and their associated mutational signatures and tumor types. The embedding process as a whole assigns vector representations that are close in the embedding space to mutations with similar mutational signature contexts. Once the input data is numerically represented in the embedding space, it can be used for mathematical, statistical, or machine learning operations and analyses, such as to perform classification or predictions about unknown data.

[0048]

[0060] Thus, in some embodiments, the machine learning model trained by machine learning module 206a may be configured to utilize embedding module 206b to pre-train mutation embeddings based on input data including multiple DNA sequencing samples, where the mutations, mutation signatures, and tumor types are represented for each input sample. The pre-training process thus learns general representations that can be used to predict contextual mutation signatures of unknown target sequencing samples.

[0049]

[0061] In some embodiments, the signature prediction module 206c may include one or more algorithms for constructing and outputting predictions 222 regarding mutational signatures represented in a target DNA sequencing sample based on the inferences of a trained machine learning model of the present disclosure.

[0050]

[0062] System 200 described herein is merely an exemplary embodiment of the present invention and may actually be implemented solely in hardware, solely in software, or a combination of hardware and software. System 200 may have more or fewer components and modules than shown, may combine two or more components, or may have components in different configurations and arrangements. System 200 may also include any other components (not shown) capable of functioning as an operational computer system, such as a motherboard, data bus, power supply, network interface card, display, input devices (e.g., keyboard, pointing device, touch panel display), etc. Furthermore, the components of system 200 may be co-located or distributed, and the system may be configured to run as one or more cloud computing “instances,” “containers,” “virtual machines,” or other types of encapsulated software applications, as known in the art.

[0051]

[0063] The instructions of system 200 will now be described with reference to the flowchart of Figure 3, which illustrates the functional steps in a method 300 for training and inferring a machine learning model configured to identify dominant mutation signatures in DNA sequencing samples obtained from a subject, according to some embodiments of the present invention.

[0052]

[0064] The various steps of method 300 will be described with successive reference to exemplary system 200 shown in Figure 2. The various steps of method 300 may be performed in the order presented, or may be performed in a different order (or in parallel), so long as the necessary inputs to a step are obtained from the outputs of previous steps. Furthermore, unless otherwise specified, the steps of method 300 may be performed automatically (e.g., by system 200 of Figure 2). Furthermore, the steps of method 300 are described for illustrative purposes, and modifications to the flowchart are anticipated, if necessary or desirable.

[0053]

[0065] Method 300 begins at step 302, where system 200, under the direction of machine learning module 206a, may receive as input a plurality of DNA sequencing samples obtained from a cohort of subjects with diverse tumor types. In some embodiments, the input sequencing samples may include whole exome sequencing (WES) and / or whole genome sequencing (WGS) mutation profiles.

[0054]

[0066] Our exemplary implementation of step 302 used DNA sequencing samples obtained from one or more of the following sources: The Cancer Genome Atlas (TCGA) Project: TCGA is using genome sequencing and bioinformatics to create a catalog of cancer-causing genetic mutations. (https: / / www.cancer.gov / about-nci / organization / ccg / research / structural-genomics / tcga) The Memorial Sloan Kettering-Integrated Mutation Profiling of Actionable Cancer Targets (MSK-IMPACT): This dataset contains mutation profiling of 10,000 sequenced cancer-related samples. (https: / / www.mskcc.org / msk-impact) The Memorial Sloan Kettering-Metastatic Events and Tropisms (MSK-MET): This dataset includes samples from a pan-cancer cohort of tumor genomic and clinical outcome data from 25,000 patients (https: / / www.cbioportal.org / study / summary?id=msk_met_2021). ·The Memorial Sloan Kettering-Immune Checkpoint Inhibitor (MSK-ICI): RM Samstein, et al. Tumor mutational load predicts survival after immunotherapy across multiple cancer types. Nature Genetics. 51, 202-206 (2019). GENIE: The American Association for Cancer Research (AACR) Project GENIE, which includes over 154,000 sequenced cancer samples from 137,000 patients. (See Powering Precision Medicine through an International Consortium. Cancer Discovery. 7, 818-831 (2017))

[0055]

[0067] In step 304, under the direction of machine learning module 206a, system 200 may receive as input data relating to at least some of the input sequencing samples received in step 302 and associate these input data with the corresponding input sequencing samples. In some embodiments, the input data received in step 304 includes, for each input sequencing sample, at least one of the following: Tumor type associated with the input sequencing sample Individual mutations present in the input sequencing sample Dominant mutation signatures associated with the input sequencing samples

[0056]

[0068] In some embodiments, the input data received in step 304 includes one or more of the following additional indicators for each input sequencing sample: Details of different cancers and tumors Tumor site Metastatic sites Structural mutations Subject's medical history o Smoking history o Life status o Survival status o Physical condition. Target demographic data oGender Age

[0057]

[0069] In some embodiments, each mutation represented in the input sequencing sample can be represented by a 5-mer sequence using a conventional classification based on six substitution subtypes and the two 5' and 3' nucleotides closest to the mutation. Overall, 6 × (4 4 ) = 1,536 possible mutation classifications. A relatively small number of individual mutations is known to introduce sparsity into the pre-training process, which increases the specificity of each mutation representation and therefore improves classification performance.

[0058]

[0070] In some embodiments, the index of dominant mutational signatures for each input sequencing sample may be based on publicly available information (see L.B. Alexandrov, et al. The repertoire of mutational signatures in human cancer. Nature. 578, 94-101 (2020); A. Zehir, et al. Mutational landscape of metastatic cancer revealed from prospective clinical sequencing of 10,000 patients. Nature Medicine. 23, 703-713 (2017)).

[0059]

[0071] In some embodiments, the index of dominant mutational signatures in each input sequencing sample is based at least in part on the relative contribution of the mutational signature in the input sequencing sample, e.g., the proportion of mutations in the input sequencing sample out of the total number of mutations associated with a particular signature, since there may be multiple active mutational signatures in each input sequencing sample with different degrees of prominence or dominance.

[0060]

[0072] In one example, we assigned a mutation signature index to each input sequencing sample based on the following methodology. Mutational signature: APOBEC o Mutation: SBS2+SBS13 Relative contribution: >30% Mutation Signature: UV o Mutations: SBS7a-d and SBS38 Relative contribution: >30% Mutation signature: Tobacco o Mutation:SBS4 Relative contribution: >30% Mutation signature: HRD o Mutation: SBS3 Relative contribution: >50% Mutation Signature: Clock_SBS1 o Mutation: SBS1 Relative contribution: >40% Mutation Signature: Clock_SBS5 o Mutation:SBS5 Relative contribution: >40% Mutation Signature: POLE o Mutation: SBS10a~b Relative contribution: >20% Mutation signature: MMR o Mutations: SBS6, SBS14, SBS15, SBS20, SBS21, SBS26, SBS44 Relative contribution: >30%

[0061]

[0073] As shown, in this particular example, to annotate an input sequencing sample with one of the APOBEC, Tobacco, UV, or MMR signatures, the relative contribution of the corresponding signature must be at least 30%. To annotate a sample with the signature HRD, a relative contribution of at least 50% is required. This is because the mutation SBS3 has a featureless, "flat" profile, which can result in a high false-positive detection rate. To annotate an input sequencing sample with the signature POLE, a threshold of 20% is required. This is because POLE has a unique mutation profile with specific characteristics, usually followed by an excessive mutation or excessive mutation phenotype. Due to their ubiquity, to annotate an input sequencing sample with the signature Clock_SBS1 or Clock_SBS5, a relative contribution of SBS1 or SBS5, respectively, must be at least 40%, and furthermore, other signatures may not be associated with the input sequencing sample.

[0062]

[0074] In alternative embodiments, different, alternative, or fewer mutation signature classes may be used to annotate the input sample data, and different annotation methodologies may be applied to the selected mutation signature classes.

[0063]

[0075] Referring back to FIG. 3, in step 306, a data pre-processing stage may be performed to apply any one or more of data cleaning, data normalization, removal of defective data, data quality control, and / or other suitable pre-processing methods or techniques.

[0064]

[0076] For example, in some embodiments, data preprocessing may include pruning of input sequencing samples received in step 302 based on accuracy level (e.g., pruning samples with an accuracy of less than 0.7, as provided by PCAWG analysis) and / or number of mutations (e.g., pruning WES samples with fewer than 20 mutations or WGS samples with fewer than 100 mutations).

[0065]

[0077] In some embodiments, data cleaning may include removing duplicate samples received from two or more different sources. In some embodiments, data cleaning may also include removing samples associated with particular mutational signatures that are sparsely represented or may be the result of sequencing artifacts, such as SBS27 and SBS45 (SBS27 is common in AML samples, and SBS45 is common in many cancer types).

[0066]

[0078] Referring back to FIG. 3, in step 308, machine learning module 206a instructs system 200 to: (i) the input sequencing sample received in step 302; (ii) a label indicating at least the tumor type, mutation data / class, and at least one dominant mutation signature associated with each input sequencing sample; A training data set may be constructed that includes:

[0067]

[0079] In one example, the mutation signature annotation labels include the following nine signatures based on the index methodology detailed above: APOBEC UV Tobacco HRD Clock_SBS1 Clock_SBS5 POLE MMR Alternate signatures (for samples where none of the previous signatures are dominant)

[0068]

[0080] However, in alternative embodiments, mutational signature annotation may be based on a desired selected set of specific mutational signatures, including, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20 or more mutational signatures.

[0069]

[0081] In one example, input sequencing samples may be annotated or labeled with a combination of two or more dominant mutation signatures. For example, such a combination may include signatures SBS5 or SBS1 (which are ubiquitous and affect many samples) combined with another signature, such as Tobacco, UV, APOBEC, MMR, and / or POLE.

[0070]

[0082] In some embodiments, the label may further indicate one or more of the following: Tumor site Metastatic sites Structural mutations Subject's medical history o Smoking history o Life status o Survival status Physical condition Target demographic data oGender Age

[0071]

[0083] In some embodiments, at step 310, machine learning module 206a may instruct system 200 to train a machine learning model of the present disclosure using the training dataset constructed at step 308.

[0072]

[0084] In one embodiment, under the direction of embedding module 206b, system 200 may pre-learn, in a training phase 310, the embedding of mutations in the association vector space based on the training dataset constructed in step 308, where at least mutation data / class, mutation signature, and tumor type are represented for each input sample.

[0073]

[0085] In some embodiments, embedding module 206b may instruct system 200 to receive as input the training dataset constructed in step 308, where each mutation is considered a token and the associated dominant signature and tumor type are considered tags. Embedding module 206b may instruct system 200 to generate positive pair associations for each token and associated tag.

[0074]

[0086] For example, the mutation CT[C>G]AA in a sample from a bladder cancer patient with an APOBEC mutation signature has a positive pair of (CT[C>G]AA, bladder cancer) and (CT[C>G]AA, APOBEC signature). For each such positive pair, embedding module 206b may instruct system 200 to generate k negative pairs consisting of tags unrelated to the mutation. In the above example, the negative pair for the related sample would include (CT[C>G]AA, colorectal cancer) and (CT[C>G]AA, MMR signature). The negative pairs are generated using a k-negative random sampling strategy.

[0075]

[0087] In some embodiments, the pre-learning training phase operates a neural embedding architecture that includes an input layer and an output layer. Each forward pass consists of computing a dot product as follows:

[0076]

number

[0077] where A and B represent the embedding vectors of all positive pairs (i.e., embeddings of different entities for the same sample) and k negative pairs. Then, we combine these dot product scores (positive scores and negative scores) to calculate the hinge loss function as follows:

[0078]

number

[0079] In the formula, b + and b - represent positive and negative pairs, respectively. At each epoch, the loss may be optimized using any suitable algorithm, such as the Decoupled Weight Decay Adam algorithm. Hyperparameters include the length of the embedding vector (l), the number of negative samples per positive sample (k), the embedding parameter (p), and the hinge loss constant margin (u). In one implementation, the following hyperparameter values ​​were used: l = 200, k = 30, p = 0.5, u = 1, with a learning rate of 1e-3, 20 epochs, a batch size of 4,096, and a seed set of 88.

[0080]

[0088] In some embodiments, step 310 includes creating a vector embedding associated with two or more dominant mutational signatures. In such cases, the average embedding vector of the two or more individual mutational signatures generated in step 310 may be used to create a new embedding vector representing the two or more mutational signatures.

[0081]

[0089] In one example, step 310 may include creating a vector embedding associated with two or more dominant mutation signatures for each combination of signatures SBS5 and SBS1 (which are ubiquitous and affect many samples) with another signature. Thus, for example, the average embedding vector of Clock_SBS5 and Tobacco may be used to create a joint embedding vector, Tobacco_SBS5, representing the combination of Tobacco and SBS5. During inference (described in step 312 below), the joint vector embedding Tobacco_SBS5 may be compared with the mutation embedding vector of the tumor sample. If the cosine similarity between the sample embedding and the joint embedding Tobacco_SBS5 is greater than the similarity scores associated with Clock_SBS5 and Tobacco, respectively, it may be concluded that a particular sample possesses both signatures.

[0082]

[0090] In some embodiments, the technique may generate combined mutational signature embeddings in step 310 for at least one or more of the following mutational signature combinations: UV_SBS5 UV_SBS1 APOBEC_SBS5 APOBEC_SBS1 MMR_SBS5 MMR_SBS1 POLE_SBS5 POLE_SBS1 TOBACCO_SBS5 TOBACCO_SBS1

[0083]

[0091] Referring back to FIG. 3, in step 312, under the direction of signature prediction module 206c, system 200 may infer the machine learning model pre-learned and trained in step 310 on an unknown target DNA sequencing sample 220 (shown in FIG. 2) obtained from the target subject and output a prediction 222 (shown in FIG. 2) regarding the mutational signature associated with target sample 220.

[0084]

[0092] In some embodiments, the inference step 312 involves extracting a pentanucleotide mutation landscape from the sample 220, as described above. For each extracted mutation, an embedding vector is identified based on the pre-training performed in step 310. The average mutation embedding vector for the sample 220 is then compared to each of the pre-trained signature embeddings (which may include a combined embedding representing two or more individual mutation signatures, such as TOBACCO_SBS5), and a corresponding similarity score is calculated between the average mutation embedding for the sample 220 and each pre-trained mutation signature.

[0085]

[0093] In some embodiments, the similarity score is based at least in part on cosine similarity, as follows:

[0086]

number

[0087] Here, a cosine value of 1 means identity vectors.

[0088]

[0094] In some embodiments, at the direction of the signature prediction module 206c, the system 200 may select and output as the predicted mutation signature 222 associated with the sample 220 the pre-trained signature embedding that exhibits the highest similarity score with the average mutation embedding of the sample 220.

[0089]

[0095] In some embodiments, signature prediction module 206c may instruct system 200 to issue a prediction only if the highest similarity score exceeds a minimum threshold, such as 0.60. If no similarity score exceeds the minimum threshold, signature prediction module 206c may instruct system 200 not to issue a prediction.

[0090] Experimental results

[0096] As described in detail below, the inventors conducted several experiments to validate the machine learning model on a variety of validation datasets.

[0091]

[0097] The validation datasets used in the experiments included DNA sequencing samples obtained from the following public datasets, unless otherwise specified: The Cancer Genome Atlas (TCGA) Project Pan-cancer Analysis of Whole Genomes (PCAWG) Project ·The Memorial Sloan Kettering-Integrated Mutation Profiling of Actionable Cancer Targets(MSK-IMPACT) ·The Memorial Sloan Kettering-Metastatic Events and Tropisms(MSK-MET) MSK-ICI GENIE

[0092] Prediction of mutational signatures using downsampled target samples

[0098] To test the ability of our trained machine learning model to accurately predict mutational signatures from target samples with limited sequencing data, we constructed a validation dataset containing 2,880 TCGA samples in which the dominant signature accounted for at least 50% of the mutations (except for Clock_SBS5, where the minimum contribution threshold was set at 75%).

[0093]

[0099] The trained machine learning model of the present disclosure was then iteratively inferred on random samples from the validation dataset based on random downsampling of the number of mutations represented in each sample using two different methodologies: (i) Randomly select only 20% of the mutations represented in each sample. (ii) Randomly select 1 to 15 sets of mutations represented in each sample.

[0094]

[0100] The results of the test are shown in Figures 4A to 4D.

[0095]

[0101] Figure 4A shows the area under the receiver operating characteristic curve (AUROC), and Figure 4B shows the area under the precision-recall curve (AUPRC) scores for predicting the dominant signature in input samples downsampled to 20% mutations randomly selected from the validation set of 2,880 samples, as described above. Values ​​are the average of 10 replicates.

[0096]

[0102] Figures 4C and 4D show the prediction results of our machine learning model for input samples randomly downsampled to 1 to 15 mutation sets, averaging 10 replicates for each mutation set. Figure 4C shows the positive predictive value (PPV) scores for different signatures. Figure 4D shows the sensitivity, specificity, PPV, and negative predictive value (NPV) for the APOBEC signature. As expected, the predictive ability of the trained machine learning model depends on the number of mutations represented in the target sample; the higher the number of mutations, the more accurate the prediction. However, this trend differs between signatures. APOBEC and UV predictions required the fewest mutations, while the MMR signature required more mutations and achieved lower prediction scores. In the case of MMR, this is likely due to mismatches with Clock_SBS1, as both signatures share a significant C>T mutation at CpG.

[0097] Prediction of signature labels in PCAWG WGS samples

[0103] We further tested this machine learning model by performing inference on an independent dataset of 2,750 PCAWG WGS samples (see Pan-cancer analysis of whole genomes. Nature. 578, 82-93 (2020)).

[0098]

[0104] This dataset contains 2,750 samples, of which 2,117 are Clock_SBS5 dominant. We first inferred our machine learning model using pre-trained embeddings without any correction for the variants represented in the samples (see step 310 in Figure 3). Compared to the previously reported TCGA WES samples, the PCAWG WGS samples contain approximately 65 times more variants (median values ​​of 83 and 5260, respectively).

[0099]

[0105] The PCAWG WGS dataset contained 1,253 samples, of which 929 were Clock_SBS5 and only Clock_SBS1. A total of 0.2% of mutations per sample were randomly selected, with a minimum of 10 mutations selected for samples with fewer than 5,000 mutations. A trained machine learning model was inferred to predict dominant signatures for the selected mutations from each sample. Overall, the classification metrics were relatively high and largely consistent with the TCGA results reported above. Therefore, our machine learning model was able to accurately predict both samples with very low mutation burden and WGS samples with hundreds of thousands of mutations.

[0100]

[0106] The prediction results are shown in Figures 5A and 5B. Figures 5A and 5B show the prediction sensitivity, specificity, PPV, and NPV scores of the machine learning model for the benchmark validation set, which included 1,153 WGS samples and was downsampled to 0.2% mutations, averaged over 10 replicates. Results are shown for the APOBEC, Clock_SBS5, MMR, POLE, Tobacco, and UV signatures.

[0101]

[0107] Furthermore, we measured the sensitivity, specificity, PPV, and NPV for APOBEC, Clock_SBS1, Clock_SBS5, HRD, MMR, Tobacco, and UV without mutation downsampling. Our machine learning model accurately predicted PCAWG samples with high evaluation scores. Furthermore, almost all misclassifications were the result of mislabeling of Clock_SBS5-dominant samples. The reason for mislabeling may be due to the fact that SBS5 is a ubiquitous signature active to some degree in almost all cells. The PPV scores for all signatures, except for SBS1 and HRD, ranged from 91.4% to 100%. Notably, only nine samples in the PCAWG WGS were labeled as Clock_SBS1, which may explain the relatively poor performance in PCAWG prediction.

[0102] Signature prediction in targeted gene panels

[0108] We constructed four targeted gene panel validation datasets from the following data sources: MSK-IMPACT MSK-MET MSK-ICI

[0103] MSK-IMPACT Panel

[0109] We selected samples with sufficiently high mutation rates (>13.8 mut / Mb) for signature analysis using a panel of 994 MSK-IMPACT targeted genes annotated with mutational signature indices based on Zehir et al. The labels and signature distribution used for this dataset are shown in Figure 6A.

[0104]

[0110] Figure 6B shows the sensitivity, specificity, PPV, and NPV scores of the model predictions in the MSK-IMPACT samples, according to the labels shown in Figure 6A.

[0105]

[0111] The MSK-IMPACT cohort consists of approximately 9000 additional samples for which no mutational signature indicators are available. For these samples, we inferred the model as follows. (i) Examining the relationship between tissue and signature (UV in skin cancer; Tobacco in lung cancer; and APOBEC in breast and bladder cancer). (ii) compare these landscapes with the landscape of dominant signatures across the entire WGS cohort; (iii) performance is compared with two common signature fitting methods. a. Non-negative least squares (NNLS) b.Mix (see I. Sason, et al. A mixture model for signature discovery from sparse mutation data. Genome Medicine. 13 (2021), doi:10.1186 / s13073-021-00988-7)

[0106]

[0112] The model was inferred for all MSK-IMPACT samples with at least four mutations. Overall, within each cancer type, the predicted signatures actually had the largest proportion: UV in skin cancer, Tobacco in lung cancer, Clock_SBS5 across many cancers, etc.

[0107]

[0113] To compare the results of our model with those of NNLS / Mix, samples were classified into five groups based on the number of mutations in each sample: 1–3, 4–7, 8–12, 13–20, and more than 20 mutations. Our model predicted a greater number of signatures for the assumed cancer types than NNLS and also compared with Mix, except for the Tobacco signature. The results are shown in Figures 6C–6H. Importantly, our model predicted 1,538 samples as Clock_SBS5, a ubiquitous, pan-cancer signature. In contrast, Mix and NNLS predicted 0 and 5 samples as Clock_SBS5, respectively.

[0108] Prediction of signature labels from the MSK-MET cohort

[0114] We constructed a separate validation dataset using additional sequencing samples obtained from the MSK-MET cohort.

[0109]

[0115] Figure 7A is a bar graph showing the number of samples from the MSK-MET cohort by mutation number.

[0110]

[0116] The MSK-MET dataset contained 9,615 samples with at least four mutations. The machine learning model was estimated for these samples. The distribution of signature predictions across cancer types showed similar results to the MSK-IMPACT cohort and PCAWG WGS samples described above. Notably, over 75% of skin cancer samples had a UV signature; Tobacco was most common in lung cancer; APOBEC was most common in bladder, breast, head and neck, and thyroid cancers; MMR was most common in endometrial, colon, and prostate samples; and POLE was most common in endometrial and colon samples, as expected. In MSK-IMPACT, only cancer types with at least 50 samples were included.

[0111]

[0117] For samples with more than three mutations, the results are very similar to those from the MSK-IMPACT and MSK-MET cohorts described above. This strengthens the prediction results and provides a landscape of these signatures across a total of over 60,000 samples. For the two- to three-mutation and single-mutation groups, the minimum cosine threshold was increased from 60% to 80% to reduce the false positive rate. Therefore, to output a prediction, the sample mutation profile embedding must have at least 80% similarity to the pre-trained signature embedding. Here, most samples were left without predictions. However, for the two- to three-mutation groups, the dominant signature in skin cancer was still the UV signature, and the dominant signature in bladder, breast, and head and neck cancers was APOBEC. Although the false positive rate was significantly higher in the single-mutation groups, Tobacco and APOBEC signatures remained dominant in the expected cancer types.

[0112]

[0118] We constructed a separate validation dataset containing 1,661 MSK-based samples from patients receiving immunotherapy. These samples, designated MSK-ICI, were derived from both the MSK-IMPACT and MSK-MET cohorts. Because MSK-ICI patients are part of the MSK-MET cohort, we removed overlapping samples from the MSK-MET cohort.

[0113]

[0119] The results are shown in Figures 7B to 7F. These are Kaplan-Meier plots of overall survival in patients treated with ICI (MSK-ICI) or MSK-MET cohorts, stratified by specific signatures detected by this machine learning model. Figure 7B shows the results for the melanoma, UV signature, and MSK-ICI cohort. Figure 7C shows the results for the melanoma, UV signature, and MSK-MET cohort. Figure 7D shows the results for the head and neck cancer, APOBEC signature, and MSK-ICI cohort. Figure 7E shows the results for the NSCLC, Tobacco signature, and MSK-ICI cohort. Figure 7C shows the results for the NSCLC, Tobacco signature, and MSK-MET cohort.

[0114]

[0120] Figures 8A and 8B show multivariate survival analyses of melanoma, UV signature, MSK-ICI (Figure 8A) and MSK-MET (Figure 8B).

[0115]

[0121] Among melanomas, 232 samples harbored at least four mutations. Samples predicted as UV-positive by this model had favorable overall survival (OS) rates, both univariately (p = 0.00038, Cox proportional hazards) and multivariately (hazard ratio 0.46, CI 0.29-0.75, p = 0.002), regardless of tumor mutational burden (TMB), primary / metastatic status, and ICI (immune checkpoint inhibitor) type (Figures 7B and 8A).

[0116]

[0122] Importantly, in MSK-MET samples that were not part of MSK-ICI and had no treatment information, UV-positive samples were still strongly associated with a favorable prognosis in both univariate (p<0.0001) and multivariate analyses (hazard ratio 0.52, CI 0.36-0.74, p<0.001) (Figures 7C and 8B). Furthermore, UV predictions from our model better discriminated survival than did annotations by Zehir et al.

[0117]

[0123] For MSK-ICI, APOBEC-predicted head and neck cancer samples had better survival than non-APOBEC samples in univariate and multivariate analyses (p=0.0015 and p=0.004, respectively) (Figure 7D).

[0118]

[0124] MSK-MET did not show significant differences in survival for APOBEC samples in head and neck cancer. This may be due to different treatment regimens, different subtype distributions, or both. Finally, in both MSK-ICI and MSK-MET, Tobacco signature-positive samples were also associated with a favorable prognosis in NSCLC in both univariate (p = 0.0026, p = 0.0049) and multivariate (p = 0.014, p = 0.03) analyses (Figures 7E and 7F). Because the PPV of Tobacco versus UV and APOBEC in NSCLC is low, only samples with at least 12 mutations were considered, rather than the four previously described.

[0119]

[0125] Figures 9A and 9B show a comparison of p-values ​​(Figure 9A) and risk values ​​(Figure 9B) between the results of the multivariate survival analyses shown in Figures 7A to 7F and 8A and 8B and the results of annotating the signatures with a combination of two mutation signatures. Improvements are observed using the dual-labeling method, except in the case of APOBEC.

[0120]

[0126] The signature landscape established based on these studies allows for analysis to detect associations between specific signatures and cancer genes or hotspot mutations. Many such associations have been identified. For example, Clock_SBS5 samples are associated with mutations in TP53, EGFR, KRAS, and others. All-cancer overall survival analyses based on Clock_SBS5 classification demonstrated an association with poor prognosis across all cancer types across independent cohorts, independent of age where available. Because TP53 and KRAS are considered negative prognostic markers, this association was also independent of mutations in these genes, thus demonstrating a more general association with survival rather than simply being an indicator of age or a specific mutated gene. To further explore the association of gene signatures, binomial tests were performed to detect positive and negative associations using Clock_SBS5 as an example. Different approaches were used to assess the association of hotspot mutations. Examples include FGFR3 with APOBEC in bladder cancer. S249C (81 / 107 and 135 / 224 in MSK-MET and GENIE, respectively) and EGFR with Clock_SBS5 in non-small cell lung cancer (NSCLC). T790M (24 / 45 and 121 / 190 in MSK-MET and GENIE, respectively). A positive association with Clock_SBS1 was also observed. As expected, Tobacco is a frequent signature in NSCLC, but EGFR 790M Only 1 / 46 (2.2%) and 2 / 310 (0.6%) of the samples with mutations had the Tobacco signature, which may explain the role of EGFR in NSCLC. L858R Mutations were also observed (all associations P<0.05, exact binomial test considering signature proportions).

[0121] Mutation Interpretability and Signature Numerical Representation

[0127] One of the challenges in implementing machine learning models in healthcare is that many models lack interpretability, leading to them being considered "black boxes." Model interpretability not only opens up new avenues for utilizing these models, but also provides the opportunity to derive additional insights rather than predictions per se.

[0122]

[0128] This example demonstrates that the model's pretrained representations indeed capture the expected information, providing insight into the relationship between expanded contextual mutations (1,536 mutation classes) and signatures. Using UV signatures as an example, the model confirmed that the cosine similarity between each of the 1,536 possible mutations actually represents what would be expected based on extensive characterization of UV-related signatures. Briefly, UV damage is primarily characterized by C>T mutations in the trinucleotide contexts TCA, TCT, TCC, CCC, CCG, and CCT. These mutation class embeddings indeed share high cosine similarity with the UV signature (see Figures 11A and 11B). To a lesser extent, several T>A, T>C, and T>G mutations can be introduced by UV damage to DNA (SBS7c-d) (2). The embeddings of these mutations, such as T>A at TTT in SBS7c, also share similarity with the UV signature embedding (Figure 11A). Because a pentanucleotide context was used, there are 16 possibilities for each trinucleotide. Interestingly, not all 16 options result in similar embeddings. For example, in one of the key mutations in SBS7a, C>T at TCA, only one subset of mutations has high similarity to the UV signature, while the others do not suggest UV embedding (Figure 11B). This is also true for T>A at TTT (Figure 11A) and many other mutations in other signatures, such as C>T at TCA in APOBEC and C>A at TCA in POLE. Collectively, these findings not only allow for the interpretation of our model's results but also provide insight into the importance of pentanucleotides in signature formation. Interestingly, the higher specificity of each mutation using 1,536 classifications may explain the predictive power of our model with only a few mutations.

[0123] Clinical use example: Non-small cell lung cancer (NSCLC)

[0129] Figures 10A to 10E show survival rates for NSCLC patients.

[0124]

[0130] Figure 10A shows stratification of patients treated with any of the EGFR-TKIs whose samples correlate with the APOBEC mutation signature prediction by drug-specific progression-free survival (left panel) and overall survival (right panel). Figure 10B shows stratification of patients treated with erlotinib whose samples correlate with the APOBEC mutation signature prediction by drug-specific progression-free survival (left panel) and overall survival (right panel). Figure 10B shows stratification of patients treated with osimertinib whose samples correlate with the APOBEC mutation signature prediction by drug-specific progression-free survival (left panel) and overall survival (right panel). Figure 10D shows drug-specific progression-free survival analysis of patients treated with ICIs whose samples correlate with the TOBACCO mutation signature prediction (left panel) or smoking history (right panel). FIG. 10E shows a cohort-level overall survival analysis for patients whose samples are associated with Clock_SBS5.

[0125]

[0131] The NSCLC GENIE-BPC data includes available genomic data (different target gene panels), detailed treatment history, treatment-specific progression-free survival (PFS), overall survival (OS), and additional refined clinical data for 1,846 NSCLC patients from four institutions (see Lavery, JA et al. A Scalable Quality Assurance Process for Curating Oncology Electronic Health Records: The Project GENIE Biopharma Collaborative Approach. JCO Clin Cancer Inform (2022) doi:10.1200 / CCI.21.00105).

[0126]

[0132] The machine learning model was inferred on this dataset and was able to predict mutational signatures across the entire GENIE-BPC cohort, with the following predictions: Tobacco (49), Clock_SBS5 (30), APOBEC (13), Clock_SBS1 (3), MMR (1), and none for POLE, HRD, or UV.

[0127]

[0133] This model found that the APOBEC signature was a robust and powerful predictive marker for patient resistance to EGFR-TKIs, such as erlotinib and osimertinib (Figures 10A-10C). For all EGFR-TKIs combined (i.e., erlotinib, osimertinib, and afatinib), APOBEC-positive patients had significantly shorter treatment-specific PFS than APOBEC-negative patients (P = 0.00025, Cox proportional hazards). This is a predictive rather than prognostic biomarker, and OS in these patients is APOBEC-independent. Therefore, this is an indicator of APOBEC-mediated resistance to EGFR-TKIs. This was also shown to be true when erlotinib and osimertinib were examined separately (PFS: P<0.05; OS: P>0.05) (Figures 10B and 10C), although the number of patients in afatinib was too small (n=27, including 4 APOBEC-positive patients) to examine it alone. Other treatment arms (i.e., ICI, chemotherapy, chemotherapy plus bevacizumab, and ALK / ROS1 inhibitors) did not demonstrate APOBEC-mediated resistance.

[0128]

[0134] The model revealed that the Tobacco signature was a positive predictive marker for ICI response for both treatment-specific PFS and OS, even excluding MSK patients who were included in the previous OS analysis (P<0.05, Figure 10D, left panel). Clinical annotation of smoking history, on the other hand, was insufficient to replicate this association (Figure 10D, right panel). Thus, the model better predicts Tobacco-associated mutagenesis processes than clinical annotation of smoking, and is a robust positive prognostic marker for ICI treatment.

[0129]

[0135] Finally, from a cohort perspective, Clock_SBS5 prediction was shown to be a prognostic factor independent of age and stage (hazard ratio 1.4, 1.13–1.6 CI, Figure 10E). Excluding MSK-based patients, this association was statistically significant (hazard ratio 1.3, 1.05–1.7 CI).

[0130]

[0136] The present invention may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions for causing a processor to perform aspects of the present invention.

[0131]

[0137] A computer-readable storage medium may be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or a suitable combination thereof. More specific, non-limiting examples of computer-readable storage media include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a device having instructions recorded thereon and mechanically encoded thereon, and a suitable combination thereof. As used herein, a computer-readable storage medium is not to be construed as being a transitory signal per se, such as a freely propagating electromagnetic wave such as an electric wave, an electromagnetic wave propagating through a transmission medium such as a waveguide (e.g., a light pulse through a fiber optic cable), or an electrical signal sent through a wire. Rather, the computer-readable storage medium is a non-transitory (ie, non-volatile) medium.

[0132]

[0138] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, fiber optic transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage on a computer-readable storage medium within the respective computing / processing device.

[0133]

[0139] Computer-readable program instructions for carrying out operations of the present invention may be either assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-configuration data, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as the "C" programming language. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network, such as a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, to carry out aspects of the present invention, electronic circuitry, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may execute the computer-readable program instructions by utilizing state information from the computer-readable program instructions to personalize the electronic circuitry. In some embodiments, an electronic circuit such as an application specific integrated circuit (ASIC) may already have computer readable program instructions embedded in it at the time of manufacture, such that the ASIC is configured to execute these instructions without programming.

[0134]

[0140] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0135]

[0141] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or another programmable data processing apparatus to manufacture a machine. Thus, the instructions, executed by the processor of the computer or another programmable data processing apparatus, create means for performing the functions / acts identified in the flowchart and / or block diagram blocks. These computer-readable program instructions may be stored on a computer-readable storage medium that can instruct a computer, programmable data processing apparatus, and / or other device to function in a particular manner. Thus, a computer-readable storage medium having instructions stored thereon includes an article of manufacture containing instructions that implement aspects of the functions / acts identified in the flowchart and / or block diagram blocks.

[0136]

[0142] The computer-readable program instructions may be loaded into a computer, another programmable data processing apparatus, or another device to cause a series of operational steps to be performed on the computer, another programmable apparatus, or another device to create a computer-executed process. Thus, the instructions executing on the computer, another programmable apparatus, or another device perform the functions / acts identified in the flowchart and / or block diagram blocks.

[0137]

[0143] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of the systems, methods, and computer program products described in various embodiments of the present invention. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing a specific logical function. It should also be noted that each block in the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, can be implemented by a dedicated hardware-based system that performs a specific function or operation or executes a combination of dedicated hardware and computer instructions.

[0138]

[0144] As used in this specification and claims, the terms "substantially," "essentially," and variations thereof, when describing a numerical value, each mean a deviation from that value of up to 20% (i.e., ±20%). Similarly, when such terms describe a range of numerical values, they mean a range that is up to 20% broader (10% above and 10% below the stated range).

[0139]

[0145] Any stated numerical range herein should be considered to specifically disclose all possible subranges and individual numerical values ​​within that range. Accordingly, each such subrange and individual numerical value constitutes an embodiment of the invention. This applies regardless of the broadness of the range. For example, reciting an integer range of 1 to 6 should be considered to specifically disclose subranges such as 1 to 3, 1 to 4, 1 to 5, 2 to 4, 2 to 6, 3 to 6, etc., and individual numbers within that range, e.g., 1, 4, and 6. Similarly, reciting a fractional range, e.g., 0.6 to 1.1, 0.6 to 0.9, 0.7 to 1.1, 0.9 to 1, 0.8 to 0.9, 0.6 to 1.1, and 1 to 1.1, and individual numbers within that range, e.g., 0.7, 1, and 1.1.

[0140]

[0146] While various embodiments of the present invention have been described for purposes of illustration, this description is not intended to be exhaustive or to be limiting to the illustrative embodiments. Many variations and modifications will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein has been selected to best explain the principles of the embodiments, practical applications or technical improvements found in the market, or to enable those skilled in the art to understand the embodiments disclosed herein.

[0141]

[0147] In the specification and claims of this application, the words "comprise," "include," and "have," and their conjugations, respectively, do not necessarily limit the elements of a list with which the word may be associated.

[0142]

[0148] In the event of a conflict between this specification and a document incorporated by reference or otherwise relied upon, this specification shall control.

Claims

1. 1. A computer-implemented method comprising: receiving as input a plurality of DNA sequencing samples corresponding to a cohort of subjects having different tumor types; annotating each of the DNA sequencing samples to indicate (a) a tumor type associated with the DNA sequencing sample, (b) one or more mutations represented in the DNA sequencing sample, and (c) at least one dominant mutation signature associated with the DNA sequencing sample; During the training stage, (i) the DNA sequencing sample; (ii) a label indicating said annotation; and training a machine learning model using a training dataset comprising: In an inference step, applying the trained machine learning model to a target DNA sequencing sample from a target subject to predict a dominant mutation signature associated with the target DNA sequencing sample; A method comprising:

2. 2. The computer-implemented method of claim 1, wherein the training step includes pre-learning an embedding in an n-dimensional vector space by the machine learning model for each of the dominant mutation signatures and the mutations represented in the DNA sequencing samples.

3. the inferring step: (i) creating a vector embedding within the n-dimensional vector space for each mutation represented in the target DNA sequencing sample; (ii) calculating a similarity score between (a) the average mutation embedding vector of all the vector embeddings generated for each mutation represented in the target DNA sequencing sample, and (b) each of the dominant mutation signature embeddings pre-trained in the training phase; (iii) selecting the pre-trained dominant mutation signature embedding that exhibits the highest similarity score as the predicted dominant mutation signature associated with the target DNA sequencing sample; 3. The computer-implemented method of claim 2, comprising:

4. The computer-implemented method of claim 3 , wherein the similarity score is based on a cosine similarity calculation.

5. 5. The computer-implemented method of claim 1, wherein the target DNA sequencing sample contains 1 to 15 mutations.

6. 6. The computer-implemented method of claim 1, wherein the annotation further indicates, for at least some of the DNA sequencing samples, at least one of the following annotation categories: tumor site, metastatic site, structural mutation, smoking history of the corresponding subject, vital status of the corresponding subject, vital status of the corresponding subject, gender of the corresponding subject, or age of the corresponding subject.

7. 7. The computer-implemented method of claim 1, wherein the annotations for the mutations represented in the DNA sequencing samples are represented by 5-mer sequences.

8. 8. The computer-implemented method of claim 1, wherein for at least some of the DNA sequencing samples, the annotation represents a combination of two dominant mutation signatures.

9. 9. The computer-implemented method of claim 8, wherein the training step further comprises pre-training a joint embedding for each of the combinations of two dominant mutation signatures.

10. the inferring step: (i) creating a vector embedding within the n-dimensional vector space for each mutation represented in the target DNA sequencing sample; (ii) calculating a similarity score between (a) the average mutation embedding vector of all the vector embeddings generated for each mutation represented in the target DNA sequencing sample, and (b) each of the pre-trained dominant mutation signature embeddings and the combined embedding; (iii) selecting the (a) pre-trained dominant mutation signature embedding or (b) combined embedding that exhibits the highest similarity score as a predicted dominant mutation signature associated with the target DNA sequencing sample; 10. The computer-implemented method of claim 9, comprising:

11. at least one hardware processor; a non-transitory computer-readable storage medium having program instructions stored thereon, the program instructions executable by the at least one hardware processor: receiving as input a plurality of DNA sequencing samples corresponding to a cohort of subjects having different tumor types; annotating each of the DNA sequencing samples to indicate (a) a tumor type associated with the DNA sequencing sample, (b) one or more mutations represented in the DNA sequencing sample, and (c) at least one dominant mutation signature associated with the DNA sequencing sample; During the training stage, (i) the DNA sequencing sample; (ii) a label indicating said annotation; and training a machine learning model using a training dataset comprising: In an inference stage, the system applies the trained machine learning model to a target DNA sequencing sample from a target subject to predict a dominant mutation signature associated with the target DNA sequencing sample.

12. 12. The system of claim 11, wherein the training step includes pre-learning an embedding in an n-dimensional vector space by the machine learning model for each of the dominant mutation signatures and the mutations represented in the DNA sequencing samples.

13. the inferring step: (i) creating a vector embedding within the n-dimensional vector space for each mutation represented in the target DNA sequencing sample; (ii) calculating a similarity score between (a) the average mutation embedding vector of all the vector embeddings generated for each mutation represented in the target DNA sequencing sample, and (b) each of the dominant mutation signature embeddings pre-trained in the training phase; (iii) selecting the pre-trained dominant mutation signature embedding that exhibits the highest similarity score as the predicted dominant mutation signature associated with the target DNA sequencing sample; The system of claim 12 , comprising:

14. The system of claim 13 , wherein the similarity score is based on a cosine similarity calculation.

15. The system of any one of claims 11 to 14, wherein the target DNA sequencing sample contains 1 to 15 mutations.

16. 16. The system of any one of claims 11 to 15, wherein the annotation further indicates, for at least some of the DNA sequencing samples, at least one of the following annotation categories: tumor site, metastatic site, structural mutation, smoking history of the corresponding subject, vital status of the corresponding subject, vital status of the corresponding subject, gender of the corresponding subject, or age of the corresponding subject.

17. 17. The system of any one of claims 11 to 16, wherein the annotations for the mutations represented in the DNA sequencing samples are represented by 5-mer sequences.

18. 18. The system of any one of claims 11 to 17, wherein for at least some of the DNA sequencing samples, the annotation represents a combination of two dominant mutation signatures.

19. 20. The system of claim 18, wherein the training step further comprises pre-training a joint embedding for each of the combinations of two dominant mutation signatures.

20. the inferring step: (i) creating a vector embedding within the n-dimensional vector space for each mutation represented in the target DNA sequencing sample; (ii) calculating a similarity score between (a) the average mutation embedding vector of all the vector embeddings generated for each mutation represented in the target DNA sequencing sample, and (b) each of the pre-trained dominant mutation signature embeddings and the combined embedding; (iii) selecting the (a) pre-trained dominant mutation signature embedding or (b) combined embedding that exhibits the highest similarity score as a predicted dominant mutation signature associated with the target DNA sequencing sample; 20. The system of claim 19, comprising:

21. 1. A computer program product including a non-transitory computer-readable storage medium having program instructions embodied thereon, The program instructions executable by at least one hardware processor: receiving as input a plurality of DNA sequencing samples corresponding to a cohort of subjects having different tumor types; annotating each of the DNA sequencing samples to indicate (a) a tumor type associated with the DNA sequencing sample, (b) one or more mutations represented in the DNA sequencing sample, and (c) at least one dominant mutation signature associated with the DNA sequencing sample; During the training stage, (i) the DNA sequencing sample; (ii) a label indicating said annotation; and training a machine learning model using a training dataset comprising: a computer program product, wherein in an inference step, the trained machine learning model is applied to a target DNA sequencing sample from a target subject to predict a dominant mutation signature associated with the target DNA sequencing sample.

22. 22. The computer program product of claim 21 , wherein the training step comprises pre-training an embedding in an n-dimensional vector space by the machine learning model for each of the dominant mutation signatures and the mutations represented in the DNA sequencing samples.

23. the inferring step: (i) creating a vector embedding within the n-dimensional vector space for each mutation represented in the target DNA sequencing sample; (ii) calculating a similarity score between (a) the average mutation embedding vector of all the vector embeddings generated for each mutation represented in the target DNA sequencing sample, and (b) each of the dominant mutation signature embeddings pre-trained in the training phase; (iii) selecting the pre-trained dominant mutation signature embedding that exhibits the highest similarity score as the predicted dominant mutation signature associated with the target DNA sequencing sample; 23. The computer program product of claim 22, comprising:

24. 24. The computer program product of claim 23, wherein the similarity score is based on a cosine similarity calculation.

25. 25. The computer program product of any one of claims 21 to 24, wherein the targeted DNA sequencing sample contains 1 to 15 mutations.

26. 26. The computer program product of any one of claims 21 to 25, wherein the annotations further indicate, for at least some of the DNA sequencing samples, at least one of the following annotation categories: tumor site, metastatic site, structural mutation, smoking history of the corresponding subject, vital status of the corresponding subject, vital status of the corresponding subject, gender of the corresponding subject, or age of the corresponding subject.

27. 27. The computer program product of any one of claims 21 to 26, wherein the annotations for the mutations represented in the DNA sequencing samples are represented by 5-mer sequences.

28. 28. The computer program product of any one of claims 21 to 27, wherein, for at least some of the DNA sequencing samples, the annotation represents a combination of two dominant mutation signatures.

29. 30. The computer program product of claim 28, wherein the training step further comprises pre-training a joint embedding for each of the combinations of two dominant mutation signatures.

30. the inferring step: (i) creating a vector embedding within the n-dimensional vector space for each mutation represented in the target DNA sequencing sample; (ii) calculating a similarity score between (a) the average mutation embedding vector of all the vector embeddings generated for each mutation represented in the target DNA sequencing sample, and (b) each of the pre-trained dominant mutation signature embeddings and the combined embedding; (iii) selecting the (a) pre-trained dominant mutation signature embedding or (b) combined embedding that exhibits the highest similarity score as a predicted dominant mutation signature associated with the target DNA sequencing sample; 30. The computer program product of claim 29, comprising:

31. 1. A method of treating cancer in a patient in need thereof, comprising: obtaining a DNA sequencing sample from the patient; applying the computer-implemented method of any one of claims 1 to 10 to predict dominant mutation signatures associated with said DNA sequencing samples; administering a cancer treatment to the patient based at least in part on the predicted dominant mutation signature; A method comprising:

32. 1. A method for predicting survival in a cancer patient, comprising: obtaining a DNA sequencing sample from the patient; applying the computer-implemented method of any one of claims 1 to 10 to predict dominant mutation signatures associated with said DNA sequencing samples; predicting survival of the patient based at least in part on the predicted dominant mutation signature; A method comprising:

33. 1. A method for predicting immunotherapy response in a cancer patient, comprising: obtaining a DNA sequencing sample from the patient; applying the computer-implemented method of any one of claims 1 to 10 to predict dominant mutation signatures associated with said DNA sequencing samples; predicting an immunotherapy response of the patient based at least in part on the predicted dominant mutation signature; A method comprising:

34. 1. A method of treating cancer in a patient in need thereof, comprising: obtaining a DNA sequencing sample from the patient; applying to the DNA sequencing samples a machine learning model trained to predict dominant mutation signatures associated with the DNA sequencing samples; administering a cancer treatment to the patient based at least in part on the dominant mutation signature in the DNA sequencing sample predicted by said applying; A method comprising:

35. 1. A method for predicting survival in a cancer patient, comprising: obtaining a DNA sequencing sample from the patient; applying to the DNA sequencing samples a machine learning model trained to predict dominant mutation signatures associated with the DNA sequencing samples; predicting survival of the patient based at least in part on the dominant mutation signature in the DNA sequencing sample predicted by said applying; A method comprising:

36. 1. A method for predicting immunotherapy response in a cancer patient, comprising: obtaining a DNA sequencing sample from the patient; applying to the DNA sequencing samples a machine learning model trained to predict dominant mutation signatures associated with the DNA sequencing samples; predicting an immunotherapy response of the patient based at least in part on the dominant mutation signature in the DNA sequencing sample predicted by said applying; A method comprising: