Personalized cancer vaccine composition with enhanced clonal coverage
The machine learning-based CloneSig+ model addresses the challenge of determining tumor clonal composition by optimizing cancer vaccine compositions to include diverse mutations, ensuring comprehensive tumor coverage and improved vaccine efficacy.
Patent Information
- Application Number
- PCT/IB2025/052664
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-22
- Filing Date
- 2025-03-13
- Publication Date
- 2026-01-29
AI Technical Summary
Current methods are unable to accurately determine the clonal composition of tumors, which is crucial for ensuring comprehensive coverage by cancer vaccines, as they either lack precision in bulk sequencing or are limited to single nucleotide variants, neglecting more immunogenic complex variants.
A machine learning-based approach using a probabilistic graphical model, such as CloneSig+, that infers tumor clones from bulk sequencing data, incorporating a broader range of mutations including SNVs, indels, RNA variants, and gene fusions, to optimize cancer vaccine compositions for enhanced clonal coverage.
This method ensures that all tumor clones are covered by the vaccine, improving the efficacy of personalized cancer vaccines by identifying more immunogenic targets and enhancing computational accuracy and efficiency.
Smart Images

Figure IB2025052664_29012026_PF_FP_ABST
Abstract
Description
PERSONALIZED CANCER VACCINE COMPOSITION WITH ENHANCED CLONAL COVERAGECROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application claims benefit to European Patent Application No. EP 24190023.2, filed on July 22, 2024, which is hereby incorporated by reference herein.FIELD
[0002] The present disclosure relates to Artificial Intelligence (Al) and machine learning (ML), and in particular to a method, system, data structure, computer program product and computer-readable medium for evaluating and enhancing cancer vaccines.BACKGROUND
[0003] Tumors are heterogeneously composed of cells that evolved from an individual cell that acquires a cancer-enabling alteration or mutation. The cancer-enabling alteration or mutation allows the cells to multiply uncontrollably. In principle, each new cell that is generated by cell division shares the same genetic makeup as the original cell. A set of cells with identical genetic makeup, passed on through cell division, is called a clone. As tumors grow, other mutations can occur, giving rise to new clones. Clones can be identified based on their genetic differences. The identification of different groups of clones is important to ensure that personalized neoantigen vaccines include vaccine elements to cover all related clones associated with a tumor.SUMMARY
[0004] In an embodiment, the present disclosure provides a computer-implemented method for machine learning -based design of a vaccine composition for cancer treatment. The method includes determining, based on mutational signatures and using a probabilistic graphical model, a set of clones associated with a collection of patient-specific tumor cells, the mutational signatures having been generated from bulk sequencing data based on a plurality of mutations occurring in tumor cells. Each clone is associated with a subset of the mutations, each clone is predicted to occur at a corresponding frequency, and genetic features are assigned to each clone based on the corresponding frequency. The vaccine composition associated with the patientspecific tumor cells is generated based on the determined set of clones. The method can be applied to use cases in medical Al and healthcare, such as for vaccine design and production, and to support or optimize decision making.BRIEF DESCRIPTION OF THE DRAWINGS
[0005] Embodiments of the present disclosure will be described in even greater detail below based on the exemplary figures. The present disclosure is not limited to the exemplaryembodiments. All features described and / or illustrated herein can be used alone or combined in different combinations in embodiments of the present disclosure. The features and advantages of various embodiments of the present disclosure will become apparent by reading the following detailed description with reference to the attached drawings which illustrate the following:
[0006] FIG. 1 schematically illustrates a method and system according to an embodiment of the present disclosure for a training and prediction procedure;
[0007] FIG. 2 schematically illustrates a structure for a model according to an embodiment of the present disclosure; and
[0008] FIG. 3 is a block diagram of an exemplary processing system, which can be configured to perform any and all operations disclosed herein.DETAILED DESCRIPTION
[0009] Embodiments of the present disclosure provide a method, system, computer-readable medium and computer program product that identify all possible clones of cancer cells associated with a tumor, and thereby use the clonal information associated with the tumor to evaluate existing compositions of cancer vaccines with regard to clonal coverage. This can be especially advantageous for improving cancer vaccines by enhancing their composition to ensure all possible clones of the tumor are covered by the cancer vaccine.
[0010] Tumor cells (e.g., cancer cells) usually evolve from an individual cell that acquires a cancer-enabling alteration or mutation, which leads to uncontrolled cell growth and division, resistance to cell death, and the ability to invade surrounding tissues. In principle, each new cell that is generated by cell division shares the same genetic makeup as the original cell. A set of cells with identical genetic makeup, passed on through cell division, is called a clone. Over the course of time, daughter cells acquire mutations that can be further passed on to respective daughter cells, creating new clones sharing the same sets of mutations.
[0011] The vast majority of mutations found in cancer cells are so-called single nucleotide variants (SNVs), which include changes of individual bases at one position. There are also more complex variants like insertions and deletions (indels), chromosomal aberrations like copy number variations (CNVs), gene fusions or RNA variants. Tumor-specific mutations are harnessed by immunotherapies, for example personalized cancer vaccines, to precisely target the tumor cells. In some embodiments, therapy using patient-individual tumor-specific mutations covers the whole tumor, for example, all of the clones present.
[0012] However, it is currently impossible to determine the clonal composition of a tumor with wet lab methods. Bulk sequencing, which generates mutation information from all cells as a batch of the sample taken gives precise information on the individual mutation but does not contain information about which group of tumor cells share which mutations. Single cellsequencing returns mutation information at the individual cell level. But this method is unable to capture nearly as many mutations occurring in the tumor cells, which is important to identify the most immunogenic targets.
[0013] Tumor clonality has also been investigated by computational scientists, who came up with probabilistic solutions to infer clone number and clone assignment of mutations from bulk sequencing data. However, these methods aim at a general understanding of clonal composition of tumors and therefore take only SNVs and copy number variation into account (see e.g., Abecassis, Judith, Fabien Reyal, and Jean-Philippe Vert. 2021. “CloneSig Can Jointly Infer Intra-Tumor Heterogeneity and Mutational Signature Activity in Bulk Tumor Sequencing Data;” Nature Communications 12 (1): 5352, which is hereby incorporated by reference herein; and see e.g., Xiao, Yao, Xueqing Wang, Hongjiu Zhang, Peter J. Ulintz, Hongyang Li, and Yuanfang Guan. 2020; “FastClone Is a Probabilistic Tool for Deconvoluting Tumor Heterogeneity in Bulk-Sequencing Samples.” Nature Communications 11 (1): 4469, which is incorporated by reference herein).
[0014] In some embodiments, the present disclosure improves cancer vaccines by informing the vaccine composition with the tumor’s clonal structure to ensure all clones are covered or to evaluate existing vaccine compositions in regard to their clonal coverage. Cancer vaccines activate a patient’s T cells by administering constructs containing their tumor-specific mutations discovered from tumor bulk sequencing data and identified as immunogenic by computational and Al methods. For this purpose, inclusion of complex variants is crucial since they tend to yield more promising vaccine targets as they are less similar to self-antigens and therefore not as tolerated by the immune system as SNVs.
[0015] Embodiments of the present disclosure provide to solve the aforementioned technical challenges and improve computer functionality to perform functions that existing technology cannot perform by providing an improved approach to infer intra-tumor heterogeneity and mutational processes by using bulk sequencing data to identify immunogenic target cells within various clones of tumor cells. In addition to enhancing computer functionality to perform functions that existing technology cannot perform, including considering a broader range of individual and joint mutational signatures for SNVs and indels, and assigning other complex variants (e.g., RNA variants and gene fusions) to clones, embodiments of the present disclosure also are able to evaluate existing vaccine compositions in regard to their clonal coverage. In addition to evaluating a vaccine composition, embodiments of the present disclosure are also able to recommend enhancements to vaccine compositions of cancer vaccines with a clonal structure associated with a tumor to ensure all clones associated with the tumor are covered by the cancer vaccine.
[0016] According to a first aspect, the present disclosure provides a computer-implemented method for machine learning-based design of a vaccine composition for cancer treatment. The method includes determining, based on mutational signatures and using a probabilistic graphical model, a set of clones associated with a collection of patient-specific tumor cells, the mutational signatures having been generated from bulk sequencing data based on a plurality of mutations occurring in tumor cells, wherein each clone is associated with a subset of the mutations, wherein each clone is predicted to occur at a corresponding frequency, and wherein genetic features are assigned to each clone based on the corresponding frequency. The vaccine composition associated with the patient-specific tumor cells is generated based on the determined set of clones.
[0017] According to a second aspect, the method according to the first aspect further comprises determining a clonal coverage score for a selected clone of the set of clones, wherein the clonal coverage score is based on whether the set of mutations of a given clone is represented by vaccine components of the vaccine composition, computing an average clonal score associated with the vaccine composition based on determining a respective clonal coverage score for each clone in the set of clones, and comparing the determined average clonal score with a threshold to determine an efficacy of the vaccine composition.
[0018] According to a third aspect, the method according to any of the first or the second aspect further comprises computing the average clonal score by assigning a corresponding priority weight to each clonal coverage score for each clone in the set of clones to generate a set of weighted clonal coverage scores, and computing the average clonal score using the weighted clonal coverage scores.
[0019] According to a fourth aspect, the method according to any of the first to the third aspects further comprises receiving a ranking of a plurality of vaccine components of the vaccine composition, which is generated based on features represented by numerical values or immunogenicity predictions of each of the vaccine components, the mutations, and copy number variations to determine frequency of variants affected by chromosomal copies.
[0020] According to a fifth aspect, the method according to any of the first to the fourth aspects further comprises the genetic features that are assigned to each clone of the set of clones comprise additional variants that are not represented by mutational signature, for example ribonucleic acid (RNA) variants or gene fusions.
[0021] According to a sixth aspect, the method according to any of the first to the fifth aspects further comprises generating the vaccine composition by optimizing a plurality of vaccine components using an optimization-based or a graph-based approach.
[0022] According to a seventh aspect, the method according to any of the first to the sixth aspects further comprises determining the set of clones by: initiating a first instance of the probabilistic graphical model, wherein the first instance of the probabilistic model takes all mutations of a first mutation type of the mutations as input, the first mutation type comprising single nucleotide variants (SNVs), training the first instance of the probabilistic graphical model by optimizing a number of clones until a validation metric associated with the first instance is optimized, and determining a first collection of clones within the patient-specific tumor cells.
[0023] According to an eighth aspect, the method according to any of the first to the seventh aspects further comprises initiating a second instance of the probabilistic graphical model, wherein the second instance of the probabilistic model takes all mutations of a second mutation types of the mutations as input, the second mutation type comprising indels, training the second instance of the probabilistic graphical model based on the optimized number of clones, and determining a second collection of clones within the patient-specific tumor cells.
[0024] According to a ninth aspect, the method according to any of the first to the eighth aspects further comprises, subsequent to initiating the first and second instances of the probabilistic graphical model, using all mutation types of the mutations for determining the set of clones.
[0025] According to a tenth aspect, the method according to any of the first to the ninth aspects further comprises performing a matching process between the first collection of clones and the second collection of clones, wherein the matching function is performed based on computing a distance function between the first collection of clones and the second collection of clones to generate a matched collection of clones.
[0026] According to an eleventh aspect, the method according to any of the first to the tenth aspects further comprises deriving clone-wise properties of interest of the matched collections of clones by computing an average or a weighted average of properties of interest from the first collection clones and the second collection of clones before matching. The weighted average can be computed by assigning a first weight to a first property of interest from the first collection clones and a second weight to a second property of interest of the second collection of clones, and computing the weighted average of the properties of interest from the first collection of clones and the second collection of clones before matching based on the assigned first weight and the second weight.
[0027] According to a twelfth aspect, the method according to any of the first to the eleventh aspects further comprises prioritizing the clones for the vaccine composition based on expressions of predetermined biomarkers or predictions of clonal fitness.
[0028] According to a thirteenth aspect, the method according to any of the first to the twelfth aspects further comprises assigning the genetic features to each clone based on a number of parameters including the corresponding frequency and additional parameters including a mutation type, a number of mutated reads, a number of reference reads, a copy number, and purity, and based on probabilistic relationships between the parameters.
[0029] A fourteenth aspect of the present disclosure provides a computer system programmed for machine learning -based design of a vaccine composition for cancer treatment, the computer system comprising one or more hardware processors which, alone or in combination, are configured to provide for execution of the method according to any of the first to thirteenth aspects.
[0030] A fifteenth aspect of the present disclosure provides a tangible, non-transitory computer-readable medium for machine learning-based design of a vaccine composition for cancer treatment, the computer-readable medium having instructions thereon, which, upon being executed by one or more processors, provides for execution of the method according to any of the first to thirteenth aspects.
[0031] FIG. 1 shows an overall system architecture for training and prediction procedure, according to embodiments of the present disclosure. A running example is used herein to demonstrate the operation and improved functionality and performance of embodiments of the present disclosure. The running example is: given variant calling (mutation information) from bulk sequencing data associated with a tumor, evaluating a cancer vaccine and determining potential ways to optimize a composition of the cancer vaccine with regard to clonal coverage.
[0032] FIG. 1 shows a training and prediction procedure 100. The training and prediction procedure utilizes an input set 104. The input set 104 includes processed data from bulk sequencing data coming from a single tumor sample. The processed data primarily comprises mutation information detected by other programs. According to some embodiments, the processed data can be generated from raw data that includes tumor (and healthy matched) bulk sequencing data. In such embodiments, the raw data is processed using one or more tools to obtain the processed data as shown in input set 104. For example, variant calling tools may be used to generate output files including mutation data. In some embodiments, the processed data includes SNVs and indels 106 and copy number variations 108 and genetic features with information (e.g., RNA variants and gene fusion) 110. Additionally, the input set 104 can include a tentative vaccine composition 112. In some embodiments, the tentative vaccine composition can be represented as plurality of vaccine elements and can also include reference DNA / RNA mutations for each vaccine element. According to other embodiments of the present disclosure, the input set 104 can also include a subset of 106, 108 and 110. In someembodiments, the subset of data from the input set 104 can be selected by external heuristic algorithms.
[0033] Large datasets of either SNVs or indels or a combination of both, as shown in 106 of FIG. 1, can be used to generate mutational signatures 102 for cancer types or other meaningful groups of samples. In some embodiments, the mutational signatures 102 are computed using data in a format as shown in 106. The mutational signatures can be probability distributions over mutation types, representing the tendency of different processes or factors to generate some kind of mutations instead of others. Unlike pre-existing methods that use mutational signatures, embodiments of the present disclosure use mutational signatures that combine both types of mutations (SNVs and indels).
[0034] In some embodiments, the joint mutational signatures 102, generated based on the variant calling (mutation detection) from bulk sequencing data samples can be a probability distribution of various types of mutations occurring in cells associated with a tumor. In some embodiments, the bulk sequencing data is further processed to obtain mutation information data in variant calling format (VCF). Based on the bulk sequencing samples obtained from multiple tumors, a probability distribution of a set of pre-defined mutation types (e.g., SNVs and indels) occurring in the tumor cells can be created. The generated probability distribution of the set of pre-defined mutations can be referred to as a mutational signature 102 associated with the tumor. In some embodiments, the mutational signature 102 models the tendency of different cellular mechanism (e.g., fault DNA mismatch repair) or exogenous causes (e.g., tobacco smoke, UV light) to lead to distinct mutational patterns within the tumor cells. For these reasons, generating the mutational signature 102 is considered to be a computationally complex and technically challenging task. In some embodiments, the mutational signature 102 is referred to as a joint mutational signature in case where the probability distribution comprises multiple different mutation types, e.g. SNVs, indels or others.
[0035] As schematically illustrated in the example of FIG. 1, among the observed (input) variables that are utilized by the probabilistic model for the prediction of tumor clones are so- called mutational signatures. A mutational signature is a discrete probability distribution over a set of mutation types. The motivation behind this definition is that there are various cellular mechanisms (e.g. faulty DNA mismatch repair) or exogenous causes (e.g. tobacco smoke, UV light) that lead to distinct mutational patterns, which can be modeled by mutational signatures (see e.g., Alexandrov, L., Nik-Zainal, S., Wedge, D. et al. Signatures of mutational processes in human cancer; Nature 500, 415-421 (2013), which is hereby incorporated by reference herein).
[0036] The adaptation of the original CloneSig algorithm (see e.g., Abecassis, Judith, Fabien Reyal, and Jean-Philippe Vert. 2021; “CloneSig Can Jointly Infer Intra-Tumor Heterogeneityand Mutational Signature Activity in Bulk Tumor Sequencing Data,” Nature Communications 12 (1): 5352) for a machine learning system according to an embodiment of the present disclosure, which system is also referred to herein as a cloner system, and which entails the assignment of mutation types beyond SNVs to clones, utilized the generation of mutational signatures covering those types of mutations to be considered.
[0037] Therefore, the cloner system according to an embodiment of the present disclosure uses a novel concept of mutational signatures, encompassing multiple types of mutations (e.g. SNVs and indels). Those signatures can be generated using a publicly available tool called Sigprofder (see e.g., Islam, S. M. Ashiqul, Marcos Diaz-Gay, Yang Wu, Mark Barnes, Raviteja Vangara, Erik N. Bergstrom, Yudou He, et al. ‘Uncovering Novel Mutational Signatures by de Novo Extraction with SigProfilerExtractor’; Cell Genomics 2, no. 11 (2022): 100179, which is hereby incorporated by reference herein) using mutation type frequencies from large data sets of tumor sequencing data. Examples of data sets that can be used for that purpose are whole exome sequencing (WES) data from the TCGA Research Network, and whole genome sequencing (WGS) data from the Pan-Cancer Analysis of Whole Genomes Network. The cloner according to an embodiment of the present disclosure uses both cancer-type-specific as well as cancer- unspecific signatures, both with options for WES and WGS data.
[0038] In some embodiments, the preexisting CloneSig algorithm is modified and improved to generate a probabilistic graphical model according to embodiments of the present disclosure that is used to predict a plurality of clones occurring in tumor cells associated with a patient, which is also referred to herein as CloneSig+ 116 for implementation in a machine learning system that generates tumor-specific vaccine compositions based on identification of clones supported by mutation signatures, in particular the cloner system 114. The CloneSig+ 116, unlike the CloneSig, is able to generate the probabilistic graphical model based on large datasets including different mutation types. The CloneSig+ 116 of the cloner system 114 is configured to utilize the mutational signature 102 to assign various types of pre-defined mutations (e.g., SNVs and indels), to different clones associated with a tumor. The mutational signature 102, as described above, is a probability distribution over the set of pre-defined mutation types.
[0039] An input sample from a patient can be used together with mutational signatures 102 to train a probabilistic model to infer the number of clones in the sample, their associated mutations and their frequency. The model is trained on both SNVs and indels and it requires joint mutational signatures to work with both at the same time but can also work with both separately. After training, RNA variants, gene fusions, and any other kind of genetic feature associated with frequency information are assigned to the inferred clones, according to theirinferred frequency. The final output of the method is then a set of clones, each associated with a list of mutations.
[0040] In some embodiments of the present disclosure, the mutational signatures 102 are utilized by the cloner system 114 to train a probabilistic model to infer a number of clones and their frequency from one of the bulk sequencing data samples of the input set 104. The probabilistic model utilized in the cloner system 114 is CloneSig+ 116, which is an improvement on the existing CloneSig model providing for enhanced computational functionality, computational efficiency and improved accuracy, among other technical improvements. For example, the CloneSig+ 116 can utilize mutational signatures 102 and the patient-specific mutations (e.g., SNVs and indels) that can be present in the tumor cells, as determined from the bulk sequencing data samples of tumor cells that are part of the input set 104 to train a probabilistic model.
[0041] The probabilistic model of CloneSig is trained for each patient and a number of clones and the model that fits best is used to predict a plurality of clones within patient-specific input data. For example, the CloneSig+ 116 creates and evaluates models for one specific patient ranging from one to n clones and following an evaluation metric, the model that generated the most promising distribution of mutations among clones is selected.
[0042] In some embodiments, RNA variants, gene fusions, and any other kind of genetic feature associated with frequency information derived from the bulk sequencing data samples of tumor cells are assigned to clones associated with the patient-specific tumor, according to the corresponding inferred frequency. In such embodiments, there is no mutational signature information available for such data, and so there are no probabilistic distributions generated for such data.
[0043] The output 118 of the CloneSig+ 116 is a predicted set of clones associated with the patient-specific tumor, each associated with a subset of mutations of one patient sample. Because generating the mutational signatures requires analyzing large amounts of complex data, and assigning the patient-specific input data with RNA variants, gene fusions, and other genetic features requires complex analysis of patient data, generating the predicted set of clones by CloneSig+ 116 is considered to be a computationally complex and technically challenging task.
[0044] Using the predicted clones and their associated mutations, the vaccine composition can be optimized, by including or excluding some targets depending on the clonal coverage of the vaccine. This process could also consider clone prioritization information coming, for instance, from clonal fitness. In alternative, the clonal coverage of a given vaccine composition can be evaluated.
[0045] In some embodiments, the output 118 of the predicted set of clones associated with the tumor, and their corresponding subset of predefined mutations are used to optimize a vaccine composition, at 120, associated with the tumor.
[0046] The list of vaccine elements contains information like identifiers to link to the variants of each clone. The list can include one or multiple scoring values usually based on immunogenicity predictions of individual vaccine elements, i.e., binding to major histocompatibility complex (MHC), to prioritize elements in cases where a subset of them needs to be selected. This information can also be provided by the order of vaccine elements in the list (ranking).
[0047] According to embodiments of the present disclosure, the vaccine elements 112 of the input set 104 further comprises additional information associated with the vaccine elements 112. In some embodiments, the additional information present along with vaccine elements 112 can also be used to link each of the vaccine elements of vaccine elements 112 to one or more clones of the predicted set of clones as determined by CloneSig+ 116 in output 118. For example, the additional information associated with the vaccine elements 112 can comprise one or multiple scoring values based on immunogenicity predictions of individual vaccine elements 112, e.g., binding to MHC. In some embodiments, one or more multiple scoring values can be based on any sort of information associated with the mutations. For example, information can include values predicted with ad-hoc machine learning models or extracted from the sequencing outputs. The scoring values associated with the vaccine elements 112 can be used to prioritize those vaccine elements 112 that are more effectively recognized as being foreign. In such cases, the vaccine elements 112 that have a higher score based on the one or multiple scoring values are selected. This information can also be provided by the order of vaccine elements in the listing (e.g., ranking) of vaccine elements 112. In some cases, the ranking of vaccine elements 112 can be based on considering a sum of scoring values calculated associated with the vaccine elements 112 and ranking the vaccine elements 112 from the highest score to the lowest score.
[0048] To optimize the vaccine composition, either of the two approaches proposed in WO 2023 / 138755 Al are used, the entire disclosure of which is hereby incorporated herein by reference. According to embodiments of the disclosure, the optimization of the vaccine composition based on vaccine elements 112 is reduced to a flow problem over a bipartite graph. A similar idea is applied in this case, using mutations and clones as nodes, and assignment probabilities as one of the scoring values used to compute the edge weights. Alternatively, the binary 0 / 1 assignments can be used instead of the assignment probability as scoring value. Furthermore, clone prioritization can be incorporated in the set of scoring values. This prioritization can be derived from methods predicting clonal fitness, for example as described inUS 63 / 552,740 or be based on other information like specific biomarker / gene expression patterns or manual ranking. In some embodiments, the prioritization, for example determined by the patterns or other predefined characteristics important to the particular application at hand, is an additional score that can be applied as a weight to the clones and / or vaccine elements within the clones.
[0049] Another possibility is the use of graph-based methods taking the phylogenetic tree, i.e., the evolutional relation of the clones as input for graph-based methods aiming to cover all nodes.
[0050] Additionally, and / or alternatively, the vaccine elements 112 of the vaccine composition associated with the tumor can be evaluated, at 124 using the output 118 of the predicted set of clones.
[0051] Embodiments of the present disclosure describe evaluating how much a clone of the set of clones determined from the patient-specific tumor data is covered by a vaccine composition, given the set of mutations predicted to appear in the clone and the vaccine elements. For example, it can be checked whether at least one of its assigned mutations corresponds to a target included in the vaccine. Alternatively, the assignment probabilities can be used rather than the binary yes / no assignments. In this case, the coverage score for a clone can be the probability that at least one of its assigned mutations is included in the vaccine.
[0052] Then, to evaluate the clonal coverage of the vaccine, the average clone coverage score can be used. These might also be weighted using the prioritization information, if this is given. For example, the average clone coverage score can be computed by summing up a plurality of clone coverage scores for the clones and dividing the computed sum by the number of clones. In case the prioritization information is provided, a weighted mean of the plurality of clone coverage scores for the clones is computed by multiplying each clone average score with associated prioritization information, summing up the resultant clone average scores and dividing the obtained sum by the number of clones. The computed average clone score can then be compared to a reference threshold to determine the efficacy of the vaccine. Additionally, and / or alternatively, the vaccine score can also be compared against vaccine scores computed for other vaccine compositions, in order to determine the best vaccine composition for a particular tumor.
[0053] According to embodiments of the present disclosure, at 126, a vaccine is evaluated by computing an average clone coverage score of the vaccine. The average clone coverage score of the vaccine provides a score that determines how many mutations of the set of mutations associated with the predicted set of clones associated with the tumor are covered by the vaccine elements 112 of the vaccine. In order to compute the average clone score, the cloner system 114can determine how many of the set of mutations associated with a selected clone of the predicted set of clones of the tumor cells, as predicted in the output 118, are covered by the vaccine. The cloner system 114 can determine whether there exists a target in the vaccine elements 112 that corresponds to at least one mutation in the set of mutations associated with the selected clone of the set of clones, as predicted in the output 118, by the CloneSig+ 116. The determination can be used by the cloner system 114 to compute a clone coverage score associated with the vaccine. The clone coverage score scores a clone based on the set of its mutations being represented by the vaccine elements. For instance, the clone coverage score, as computed by the cloner system 114 determines whether at least one mutation of the set of mutations associated with the selected clone of the predicted set of clones of the tumor cells corresponds to a target included in the vaccine elements 112. The clone coverage score calculated for each clone that is predicted to occur in the tumor by CloneSig+ 116 is used to generate an average clone coverage score associated with the vaccine. In some embodiments, the average clone coverage score can be compared to a predetermined threshold to determine an efficacy of the vaccine.
[0054] In some cases, the determination of whether there exists a target in the vaccine elements 112 that corresponds to a mutation in the set of mutations associated with the set of clones, as predicted in the output 118, is performed using a probabilistic score instead of a binary yes / no assignment. In such cases, the clone coverage score can be a probability that at least one mutations of the set of mutations associated with a clone corresponds to a target included in the vaccine elements 112. The clone coverage score calculated for each clone that is predicted to occur in the tumor by Clone Sig+ 116 is used to create an average clone coverage score associated with the vaccine.
[0055] In some embodiments, the average clone score can be weighted based on provided prioritization information.
[0056] FIG. 2 shows the graph model 200, as originally published, for Clone Sig.
[0057] Existing Clone Sig model 202 is based on a probabilistic graphical model that encodes the dependencies between the input tumor data (depicted by nodes 204, 208, 214, 216, and 218) and some underlying hidden factors (depicted by nodes 206, 212, and 210). The model takes as input, for each mutation in the bulk sequencing sample data, the number of mutated and reference reads (nodes B 208 and D 214), the mutation type (node T 216) and the copy number information for the mutation locus (node C 204). It also takes as input the sample purity (node p 218). The model is then able to estimate the clone index (node U 210), the mutated copy number (node M 206) for each mutation and the signature profile (node S 212), e.g., a probability distribution over mutational signature, for each clone. Furthermore, the model 202 can be used to estimate how likely a mutation is to appear in a specific clone.
[0058] Preexisting CloneSig methods also include a validation pipeline to infer the number of clones in the sample. This trains models with increasingly more clones until a validation metric, considering both the model complexity and performance, stops improving.
[0059] According to embodiments of the present disclosure, CloneSig+ 116 can use two instances of CloneSig: one taking in input only the SNV mutations (base model in the following) and one taking in input only the indels mutations (complex model in the following). CloneSig+ 116 can also take all mutation as input at once, which requires mutational signature fdes containing both types of mutations. By using two instances of the CloneSig 202, single nucleotide variants and indels occurring in the tumor cells can be included instead of using only single nucleotide variants as is done in conventional CloneSig methods. At first, the base model is trained by optimizing the number of clones with the validation pipeline of the preexisting CloneSig method. In some embodiments, instances of CloneSig that are part of CloneSig+ 116 are unsupervised models that are not trained with pre-assigned predictions. In such embodiments, the instances of CloneSig, much like other unsupervised learning models, are trained using the data which is to be investigated for the generation of distributions. For example, the base model can be trained using the distributions of mutations among clones. In order to perform the training, the number of clones are fixed as a parameter, and the instance of CloneSig can be optimized with a model selection procedure. In some cases, the output of the CloneSig+ 116 includes assigning the mutations to the fixed number of clones. In accordance with some embodiments, the assignment can be hard assignment or a soft assignment. In a hard assignment, each mutation is assigned to exactly one clone, whereas in case of soft assignment each mutation is assigned a probability vector where the nthentry specifies the probability that the mutation belongs to clone n.
[0060] In some embodiments, the fixed number of clones is optimized by a validation pipeline associated with the CloneSig+ 116. For each selected number of clones starting from one, a new solution fit is generated. Once a validation metric associated with the validation pipeline stops increasing, the optimization process is stopped and the number of clones is fixed.
[0061] In such cases, the training procedure of the CloneSig model can be an expectationmaximization algorithm in which the number of clones is a parameter that is set before the training. After the base model is trained, the complex model is trained with a fixed number of clones, set to the one found by the base model, as well as the base model distributions as initialization.
[0062] Since the two models learn their own set of clones, both sets of clones need to be matched together to achieve a consensus set of clones with both types of mutation associated to these clones. To achieve this, CloneSig+ 116 finds a mapping from clones in one model toclones in the other model. This is done by matching clones according to the similarity of the signature profdes (e.g., the probability distribution over the respective mutational signatures). A distance function between probability distributions (e.g., the Jensen-Shannon similarity) is assumed. First the distance of all complex clones to the base clones is computed. If the closest base clone of each complex clone is different, then they are used as matchings. If this is not the case, the same thing is tried for base clones, to check whether the closest complex clone of each base clone is different. If this is also not the case, the closest pair of clones is manually matched, and they are removed from the matching algorithm. This process is repeated until all the clones are matched. As described previously, the clones are matched to attain a unified set of clones, rather than two independent sets (one from the base model and one from the complex model), because the matched clones are generated by similar processes and are likely to have similar mutational signatures. The rest of the clone related distributions are computed by combining the distributions of the matched clones, weighting each clone with the number of mutations supporting the corresponding model.
[0063] According to embodiments of the present disclosure, one instance of CloneSig+ 116 is configured to combine the outputs of the two instances as a single version of the distribution of the clones. For example, as discussed previously, the output of the two instances of CloneSig, that are part of the CloneSig+ 116, includes tabular data that assigns each mutation to one out of the fixed number of clones. In accordance with some embodiments, the assignment can be hard assignment or a soft assignment. In a hard assignment, each mutation is assigned to exactly one clone, whereas in case of soft assignment each mutation is assigned a probability vector where the nthentry specifies the probability that the mutation belongs to clone n. The output can also include clone frequency information for each clone, and information about mutational signature activity in each clone.
[0064] In order to prepare the single output, the output of the two instances of the CloneSig, as described above are combined by computing an average of the output of the two instances, based on the number of mutations considered in each instance. For example, the average of the output of the two instances is calculated on a clone-by-clone basis. This does not affect the distribution of mutations to clones, since the two models work on a disjoint set of the input mutations (SNVs and indels), and therefore does not affect the vaccine evaluation.
[0065] In some embodiments, the two instances of CloneSig can consider a different number of mutations, and therefore, a weighted average of the outputs of the two instances, based on the number of mutations covered, can be a more reliable indication of the distribution of the clones. In such cases, the number of mutations (in percentage to the total) is used as weight for the distribution of each instance of the CloneSig before being combined.
[0066] According to some alternative embodiments, once a coherent and unique set of clones are determined, the input mutations are assigned to each clone. For the SNVs and indels, the assignment is based on the probability directly inferred by the models. For instance, a mutation can be assigned to a clone if the probability of appearing in that clone is higher than a given threshold. For the frequency-based variants, the estimated clone frequency is used instead. For instance, a mutation is assigned to the clone whose frequency is most similar to the mutation frequency, taking copy number variations into account. This might lead to multiple scenarios of mutation assignment. For those cases, all downstream evaluations and calculations are done over all scenarios.
[0067] Embodiments of the present disclosure provide the following improvements and advantages over existing technology:1) Mutational signatures / joint signatures of multiple / diverse mutation types: The clonal deconvolution of bulk sequencing data that is performed by the probabilistic clonality model can only be informed by those mutation types which are covered by the mutational signatures applied in the model. Previous models using mutational signatures only considered one mutation type (typically SNVs) for the deconvolution problem. The cloner system according to embodiments of the present disclosure comprises novel mutational signatures that cover multiple mutation types, which offers two advantages: Firstly, the identification of clones is informed by a greater number of mutations, which can improve the precision of the method. Secondly, it allows the assignment of mutations other than SNVs to the identified clones. This is important because non-SNV mutations are often more immunogenic and hence of more interest in the design of the personalized cancer vaccines.2) Addition of complex variants (see step 2) c. of the method below) and linking to existing clones by frequency: The cloner system according to embodiments of the present disclosure also enables the assignment of mutations to clones beyond those mutation types covered by the mutational signatures. The information contained in the patient’s tumor sequencing data can be combined with the clonality model prediction to map complex DNA or RNA variants and features to the predicted clones.3) Generation of vaccine composition taking clonality into account: After hard or soft assignment of the tumor mutations to clones, the cloner system according to embodiments of the present disclosure can generate suggested vaccine compositions (a shortlist of tumor-specific mutations to be included in the personalized cancer vaccine) that are optimized for covering all clones of the tumor. The generated lists can then further be refined by using other tools which take into account immunogenicity, distinction from self-antigens, etc.4) Vaccine evaluation in terms of clonal coverage: Given a vaccine composition, the clonality predictions of the cloner system according to embodiments of the present disclosure can be used to evaluate the vaccine composition in terms of its coverage of the tumor clones.
[0068] The present disclosure may be implemented as a computer-implemented method, computer system (comprising one or more processors and one or more storage devices) configured to perform the computer-implemented method and / or as a computer program for performing the computer-implemented method. For example, the computer-implemented method may include one or more steps and / or operations discussed above.
[0069] Various aspects of the present disclosure relate to machine learning. In particular, the model mentioned above may be a machine learning model.
[0070] Machine learning is a branch of artificial intelligence that involves the development of algorithms and models that allow computers to learn and make predictions or decisions without being explicitly programmed. It focuses on creating systems that can improve their performance over time by learning from data.
[0071] Training a machine-learning model refers to the process of teaching the model to make accurate predictions or decisions. During training, the model is exposed to a large amount of data, which is used to adjust the model's internal parameters or weights. The model learns patterns, relationships, or rules from the training data, allowing it to generalize and make predictions on new, unseen data.
[0072] Training data is the set of examples or instances that is used to teach a machinelearning model. It is often labeled data, meaning that each example is associated with a known outcome or target value. The training data consists of both input features and the corresponding output or target variable. The model learns from this data by analyzing the patterns and relationships between the input features and the target variable. Training algorithms, such as supervised learning, semi-supervised learning, unsupervised learning or reinforcement learning may be used fortraining the machine-learning model.
[0073] Machine-learning models, such as the machine-learning model being trained in the present disclosure, are often implemented as Artificial Neural Networks (ANNs), and in particular Deep Neural Networks, Support Vector Machines, Decision Tree models, or Random Forest models.
[0074] Examples may involve or relate to computer programs, including program codes to execute one or more of the mentioned methods when the program is executed on a computer, processor, or other programmable hardware component. As a result, steps, operations, or processes from various methods described above can also be executed by computers, processors, or other programmable hardware components. Examples may additionally cover programstorage devices, such as digital data storage media, which are machine-, processor-, or computer-readable and encode and / or contain machine-executable, processor-executable, or computer-executable programs and instructions. These devices may include or be digital storage devices, magnetic storage media like magnetic disks and tapes, hard disk drives, or optically readable digital data storage media, for instance. Other examples encompass computers, processors, control units, field programmable logic arrays (FPLAs), field programmable gate arrays (FPGAs), graphics processing units (GPUs), application-specific integrated circuits (ASICs), integrated circuits (ICs), or system-on-a-chip (SoC) systems that are programmed to carry out the steps of the aforementioned methods. In simpler terms, examples may involve computer programs and storage media comprising computer programs, as well as hardware components like processors and control units, which can be programmed to execute the methods described above.
[0075] When certain aspects are mentioned in relation to a device or system, they should also be considered as descriptions of the corresponding methods. For example, a block, component, or functional aspect of the device or system may correspond to a method step or feature of the related method. Therefore, aspects described regarding a method should also be understood as depicting a corresponding element, property, or functional feature of the corresponding device or system. In simpler terms, if something is described in relation to a device or system, it can also be applied to the corresponding method, and vice versa.
[0076] The preexisting CloneSig model (see e.g., Abecassis, Judith, Fabien Reyal, and Jean- Philippe Vert. 2021. “CloneSig Can Jointly Infer Intra-Tumor Heterogeneity and Mutational Signature Activity in Bulk Tumor Sequencing Data.” Nature Communications 12 (1): 5352, which is hereby incorporated by reference herein) is a computational method that can simultaneously infer intra-tumor heterogeneity (ITH) and mutational processes within tumors from bulk sequencing data. CloneSig accounts for potential dependencies between different mutational processes at various tumor stages or regions. It employs a generative probabilistic graphical model, treating somatic mutations as arising from a mixture of clones where distinct mutational signatures are active. This approach allows for an analysis of tumor evolution and mutational processes.
[0077] Embodiments of the present disclosure improve upon the preexisting CloneSig model to generate CloneSig+ 116 (as shown in FIG. 1). The CloneSig+ 116 improves on the preexisting model by: 1) considering a broader range of mutational signatures for SNVs and indels, 2) assigning other complex variants like RNA variants and gene fusions to clones, 3) integrating information from RNA-sequencing data in addition to whole exome sequencing(WES) and whole genome sequencing (WGS) data, 4) performing vaccine element prioritization.
[0078] FastClone is another preexisting algorithm for detecting tumor heterogeneity (see e.g., Xiao, Yao, Xueqing Wang, Hongjiu Zhang, Peter J. Ulintz, Hongyang Li, and Yuanfang Guan. 2020. “FastClone Is a Probabilistic Tool for Deconvoluting Tumor Heterogeneity in Bulk-Sequencing Samples.” Nature Communications 11 (1): 4469), which is hereby incorporated by reference herein. FastClone relies on constructing a phylogenetic tree to infer tumor heterogeneity by exhaustively exploring all possible branching options. This method processes bulk DNA-sequencing data from a single tumor sample to infer the subclonal composition. Utilizing information on copy number profiles and allele frequencies, it deconvolutes subclones that exhibit independent copy number variation events within the same chromosomal regions. However, embodiments of the present disclosure provide a different computational approach compared to the approach described in FastClone.
[0079] FastClone simply covers Single Nucleotide Variants (SNVs) in the input bulk DNA- sequencing data. Unlike embodiments of the present disclosure, FastClone does not accommodate other kind of DNA or RNA variants. FastClone is also computationally expensive because it explores all possible branching possibilities of a phylogenetic tree to establish a valid tree structure. FastClone lacks modules like clone prioritization or modules for inferring vaccine composition and evaluation.
[0080] Many modifications and other embodiments of the disclosure set forth herein will come to mind the one skilled in the art to which the disclosure pertains having the benefit of the teachings presented in the foregoing description and the associated drawings. Therefore, it is to be understood that the disclosure is not to be limited to the specific embodiments disclosed and that modifications and other embodiments are intended to be included within the scope of the disclosure. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.
[0081] Embodiments of the present disclosure thus provide for general improvements to computers in machine learning systems to generating and evaluating vaccine compositions associated with a tumor. Moreover, embodiments of the present disclosure can be practically applied to use cases to effect further improvements in a number of technical fields including, but not limited to, medical and healthcare (e.g., digital medicine, personalized healthcare, drug or vaccine development, prescription, treatment, etc.). For instance, embodiments of the present disclosure can be used to operate wet lab equipment based on the determined cancer vaccine composition to test or to generate vaccines to treat cancerous tumors. Embodiments of thepresent disclosure can also be used to efficiently generate personalized treatment plans for patients based on patient-specific data in an automated manner.
[0082] In an embodiment, the present disclosure provides a method and system machine learning-assisted vaccine design comprising the following features:1) Generation and input of mutational signatures based on large cancer datasets to quantify mutations for different settings like cancer specificity and genomic region. The generation of such mutational signatures is a computationally complex task not capable of being performed in the human mind.2) Sample-Zpatient-specific input derived from DNA-sequencing samples of a tumor of a patient. The genome of each patient as well as their tumor is unique. a. SNVs and indels derived from any variant calling pipeline. In some embodiments, a variant calling pipeline is a way of determining mutations (e.g., SNVs and indels) from sequencing data. b. Copy number variations to estimate frequency of variants that are affected by chromosomal copies. c. Gene fusions, RNA variants or other genetic features including frequency estimates (genomic or transcriptomic). d. A list of vaccine elements including ranking information. In some embodiments, the vaccine elements can be received from any software and / or pipeline that predicts and / or designs vaccine elements for personalized cancer vaccines. The vaccine elements can be ranked based on additional information and scores, such as immunogenicity.3) The core modules a. Generation of a set of clones and their assigned mutation by applying CloneSig+ to identify the most likely number of clones using mutational signature information and association of complex variants with frequency information. b. Generation of a vaccine composition taking clonality information (see step 3) a.) and vaccine elements (see step 3) d.) into account, either applying an optimization-based or graph-based approach. c. Evaluation of existing vaccine composition in regards of clonal coverage by matching vaccine elements to variants of clones.4) (in some embodiments) Clone prioritization, for example derived from prediction of clonal fitness or expression of certain biomarkers in form of ranking or other numerical value. In some embodiments, the methods described herein are agnostic to how the clonal prioritization is predicted.
[0083] Referring to FIG. 3, a processing system 300 can include one or more processors 302, memory 304, one or more input / output devices 306, one or more sensors 308, one or more user interfaces 310, and one or more actuators 312. Processing system 300 can be representative of each computing system disclosed herein.
[0084] Processors 302 can include one or more distinct processors, each having one or more cores. Each of the distinct processors can have the same or different structure. Processors 302 can include one or more central processing units (CPUs), one or more graphics processing units (GPUs), circuitry (e.g., application specific integrated circuits (ASICs)), digital signal processors (DSPs), and the like. Processors 302 can be mounted to a common substrate or to multiple different substrates.
[0085] Processors 302 are configured to perform a certain function, method, or operation (e.g., are configured to provide for performance of a function, method, or operation) at least when one of the one or more of the distinct processors is capable of performing operations embodying the function, method, or operation. Processors 302 can perform operations embodying the function, method, or operation by, for example, executing code (e.g., interpreting scripts) stored on memory 304 and / or trafficking data through one or more ASICs. Processors 302, and thus processing system 300, can be configured to perform, automatically, any and all functions, methods, and operations disclosed herein. Therefore, processing system 300 can be configured to implement any of (e.g., all of) the protocols, devices, mechanisms, systems, and methods described herein.
[0086] For example, when the present disclosure states that a method or device performs task “X” (or that task “X” is performed), such a statement should be understood to disclose that processing system 300 can be configured to perform task “X”. Processing system 300 is configured to perform a function, method, or operation at least when processors 302 are configured to do the same.
[0087] Memory 304 can include volatile memory, non-volatile memory, and any other medium capable of storing data. Each of the volatile memory, non-volatile memory, and any other type of memory can include multiple different memory devices, located at multiple distinct locations and each having a different structure. Memory 304 can include remotely hosted (e.g., cloud) storage.
[0088] Examples of memory 304 include a non-transitory computer-readable media such as RAM, ROM, flash memory, EEPROM, any kind of optical storage disk such as a DVD, a Blu- Ray® disc, magnetic storage, holographic storage, a HDD, a SSD, any medium that can be used to store program code in the form of instructions or data structures, and the like. Any and all of the methods, functions, and operations described herein can be fully embodied in the form oftangible and / or non-transitory machine-readable code (e.g., interpretable scripts) saved in memory 304.
[0089] Input-output devices 306 can include any component for trafficking data such as ports, antennas (i.e., transceivers), printed conductive paths, and the like. Input-output devices 306 can enable wired communication via USB®, DisplayPort®, HDMI®, Ethernet, and the like. Input-output devices 306 can enable electronic, optical, magnetic, and holographic, communication with suitable memory 306. Input-output devices 306 can enable wireless communication via WiFi®, Bluetooth®, cellular (e.g., LTE®, CDMA®, GSM®, WiMax®, NFC®), GPS, and the like. Input-output devices 306 can include wired and / or wireless communication pathways.
[0090] Sensors 308 can capture physical measurements of environment and report the same to processors 302. User interface 310 can include displays, physical buttons, speakers, microphones, keyboards, and the like. Actuators 312 can enable processors 302 to control mechanical forces.
[0091] Processing system 300 can be distributed. For example, some components of processing system 300 can reside in a remote hosted network service (e.g., a cloud computing environment) while other components of processing system 300 can reside in a local computing system. Processing system 300 can have a modular design where certain modules include a plurality of the features / functions shown in FIG. 3. For example, I / O modules can include volatile memory and one or more processors. As another example, individual processor modules can include read-only-memory and / or local caches.
[0092] While embodiments of the disclosure have been illustrated and described in detail in the drawings and foregoing description, such illustration and description are to be considered illustrative or exemplary and not restrictive. It will be understood that changes and modifications may be made by those of ordinary skill within the scope of embodiments of the present disclosure. In particular, the present disclosure covers further embodiments with any combination of features from different embodiments described above and below. Additionally, statements made herein characterizing the invention or disclosure refer to an embodiment of the invention or disclosure and not necessarily all embodiments.
[0093] The terms used in the claims should be construed to have the broadest reasonable interpretation consistent with the foregoing description. For example, the use of the article “a” or “the” in introducing an element should not be interpreted as being exclusive of a plurality of elements. Likewise, the recitation of “or” should be interpreted as being inclusive, such that the recitation of “A or B” is not exclusive of “A and B,” unless it is clear from the context or the foregoing description that only one of A and B is intended. Further, the recitation of “at least oneof A, B and C” should be interpreted as one or more of a group of elements consisting of A, B and C, and should not be interpreted as requiring at least one of each of the listed elements A, B and C, regardless of whether A, B and C are related as categories or otherwise. Moreover, the recitation of “A, B and / or C” or “at least one of A, B or C” should be interpreted as including any singular entity from the listed elements, e.g., A, any subset from the listed elements, e.g., A and B, or the entire list of elements A, B and C.
Claims
CLAIMSWhat is claimed is:
1. A computer-implemented method for machine learning -based design of a vaccine composition for cancer treatment, the method comprising: determining, based on mutational signatures and using a probabilistic graphical model, a set of clones associated with a collection of patient-specific tumor cells, the mutational signatures having been generated from bulk sequencing data based on a plurality of mutations occurring in tumor cells, wherein each clone is associated with a subset of the mutations, wherein each clone is predicted to occur at a corresponding frequency, and wherein genetic features are assigned to each clone based on the corresponding frequency; and generating the vaccine composition associated with the patient-specific tumor cells based on the determined set of clones.
2. The method according to claim 1, further comprising: determining a clonal coverage score for a selected clone of the set of clones, wherein the clonal coverage score is based on whether the set of mutations of a given clone is represented by vaccine components of the vaccine composition; computing an average clonal score associated with the vaccine composition based on determining a respective clonal coverage score for each clone in the set of clones; and comparing the determined average clonal score with a threshold to determine an efficacy of the vaccine composition.
3. The method of according to claim 2, wherein computing the average clonal score comprises: assigning a corresponding priority weight to each clonal coverage score for each clone in the set of clones to generate a set of weighted clonal coverage scores; and computing the average clonal score using the weighted clonal coverage scores.
4. The method according to any of the preceding claims, further comprising receiving a ranking of a plurality of vaccine components of the vaccine composition, which is generated based on features represented by numerical values or immunogenicity predictions of each of the vaccine components, the mutations, and copy number variations to determine frequency of variants affected by chromosomal copies.
5. The method according to any of the preceding claims, wherein the genetic features assigned to each clone of the set of clones comprises additional variants that are not represented by a mutational signature including ribonucleic acid (RNA) variants and / or gene fusions.
6. The method according to any of the preceding claims, wherein generating the vaccine composition comprises optimizing a plurality of vaccine components using an optimization-based or a graph-based approach.
7. The method according to any of the preceding claims, wherein determining the set of clones comprises: initiating a first instance of the probabilistic graphical model, wherein the first instance of the probabilistic model takes all mutations of a first mutation type of the mutations as input, the first mutation type comprising single nucleotide variants (SNVs); training the first instance of the probabilistic graphical model by optimizing a number of clones until a validation metric associated with the first instance is optimized; and determining a first collection of clones within the patient-specific tumor cells.
8. The method according to claim 7, wherein determining the set of clones comprises: initiating a second instance of the probabilistic graphical model, wherein the second instance of the probabilistic model takes all mutations of a second mutation types of the mutations as input, the second mutation type comprising indels; training the second instance of the probabilistic graphical model based on the optimized number of clones; and determining a second collection of clones within the patient-specific tumor cells.
9. The method according to claim 8, wherein, subsequent to initiating the first and second instances of the probabilistic graphical model, all mutation types of the mutations are used for determining the set of clones.
10. The method according to claim 8, wherein determining the set of clones comprises: performing a matching process between the first collection of clones and the second collection of clones, wherein the matching function is performed based on computing a distance function between the first collection of clones and the second collection of clones to generate a matched collection of clones.
11. The method according to claim 10, further comprising deriving clone-wise properties of interest of the matched collection of clones by computing an average or a weighted average of the clone-wise properties of interest from the first collection of clones and the second collection of clones before matching.
12. The method according to any of the preceding claims, further comprising: prioritizing the clones for the vaccine composition based on expressions of predetermined biomarkers or predictions of clonal fitness.
13. The method according to claim 1, wherein the genetic features are assigned to each clone based on a number of parameters including the corresponding frequency and additional parameters including a mutation type, a number of mutated reads, a number of reference reads, a copy number, and purity, and based on probabilistic relationships between the parameters.
14. A computer system programmed for machine learning -based design of a vaccine composition for cancer treatment, the computer system comprising one or more hardware processors which, alone or in combination, are configured to provide for execution of the following steps: determining, based on the mutational signature and using a probabilistic graphical model, a set of clones associated with a collection of patient-specific tumor cells, the mutational signatures having been generated from bulk sequencing data based on a plurality of mutations occurring in tumor cells of multiple patients, wherein each clone is associated with a subset of the mutations, wherein each clone is predicted to occur at a corresponding frequency, and wherein genetic features are assigned to each clone based on the corresponding frequency; and generating the vaccine composition associated with the patient-specific tumor cells based on the determined set of clones.
15. A tangible, non-transitory computer-readable medium for machine learningbased design of a vaccine composition for cancer treatment, the computer-readable medium having instructions thereon, which, upon being executed by one or more processors, provides for execution of the following steps: determining, based on the mutational signature and using a probabilistic graphical model, a set of clones associated with a collection of patient-specific tumor cells, the mutational signatures having been generated from bulk sequencing data based on a plurality of mutations occurring in tumor cells of multiple patients, wherein each clone is associated with a subset of the mutations, wherein each clone is predicted to occur at a corresponding frequency, and wherein genetic features are assigned to each clone based on the corresponding frequency; and generating the vaccine composition associated with the patient-specific tumor cells based on the determined set of clones.
Citation Information
Patent Citations
Methods of vaccine design
WO2023138755A1
Methods and systems for determining somatic mutation clonality
WO2019109086A1
Process for preparation of neopepitope-containing vaccine agents
WO2022023521A2
Methods for optimizing tumor vaccine antigen coverage for heterogenous malignancies
WO2022197630A1
Identification of clonal neoantigens and uses thereof
WO2023194486A1